Conversation generation program, conversation generation method, and conversation generation apparatus
The system addresses the challenge of managing conversation flow in LLMs by evaluating and generating responses that align with conversation intent, enhancing natural interaction through structured evaluation and response selection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-22
Smart Images

Figure 2026085156000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a conversation generation program, a conversation generation method, and a conversation generation device.
Background Art
[0002] The use of systems for natural language question-and-answer using large language models (LLMs) is progressing. An LLM outputs an answer to a question by inputting a sentence corresponding to the question described in natural language.
[0003] For example, as an LLM is used as a virtual speaker to conduct role-play through a pseudo-conversation with a target person who is an actual human, training for conversation with the target person is carried out. There are no particular restrictions on the situations assumed for role-play, but for example, use for measures against customer harassment in a call center can be considered.
[0004] In measures against customer harassment and the like, since the trainer side that trains the target person is required to have the skill to accurately understand the customer situation, the burden on the trainer becomes large in face-to-face training. Therefore, training of the target person is carried out by having an employee of the call center as the target person to pseudo-experience customer harassment and the like through role-play with an LLM.
[0005] Also, as a conversation technique using an LLM, a technique has been proposed in which a plurality of databases and an LLM are combined and questions are processed step by step to output an appropriate answer from the LLM for complex questions. A technique has been proposed in which a guidance mechanism for customizing an LLM is used together with the implementation requirements of a task to process the task so that various tasks can be executed with a single model.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
[0007] However, while LLM can provide accurate answers to questions in a question-and-answer format, it struggles to provide answers that consider the flow of a conversation across multiple exchanges, especially in conversations where branching occurs depending on the flow. To enable LLM to provide answers that reflect the flow of the conversation, information that helps it understand the conversation flow would need to be input into the LLM. However, it is difficult to narrow down the scope of information that represents the flow of the conversation, and there is a risk that the amount of information to be input will become enormous. Preparing such information is difficult, and processing such a large amount of information may increase the processing load on the LLM.
[0008] Furthermore, while techniques that combine multiple databases and LLMs to process questions step-by-step can answer complex or ambiguous questions, they are still limited to a question-and-answer format and do not consider the flow of conversation. Additionally, while guidance mechanisms customize LLMs to suit the task, it is difficult to optimize LLM operation in processes where branching occurs in the flow of conversation. Thus, conventional technologies have made it difficult to have natural conversations because answers are returned without considering the flow of conversation.
[0009] The disclosed technology was developed in view of the above, and aims to provide a conversation generation program, conversation generation method, and conversation generation device that enable natural conversation. [Means for solving the problem]
[0010] In one embodiment of the conversation generation program, conversation generation method, and conversation generation apparatus disclosed in this application, a first process is performed to evaluate the statements of a person in a conversation about a specific topic with a predetermined situation setting, for each of the pre-set evaluation items. The computer is then instructed to perform a process of referencing information that associates the combination of evaluations for each evaluation item with intention information indicating the intention of the person the subject is talking to in the next conversation, selecting the intention information corresponding to the combination of evaluations for each evaluation item obtained in the first process, and generating a response based on the selected intention information and the situation setting information. [Effects of the Invention]
[0011] In one respect, the present invention can provide effective conversation training. [Brief explanation of the drawing]
[0012] [Figure 1] Figure 1 shows an overview of the training system. [Figure 2] Figure 2 is a block diagram of the conversation generation device. [Figure 3] Figure 3 shows an example of the evaluation results from the first LLM. [Figure 4] Figure 4 shows an example of a correspondence table that associates evaluation results with prompts. [Figure 5] Figure 5 shows an overview of the processing performed by the conversation generation device according to Example 1 in training. [Figure 6] Figure 6 is a flowchart of the training process for a subject using the conversation generation device according to Example 1. [Figure 7] Figure 7 is a hardware configuration diagram of the conversation generation device. [Modes for carrying out the invention]
[0013] Hereinafter, embodiments of the conversation generation program, conversation generation method, and conversation generation apparatus disclosed in the present application will be described in detail based on the drawings. Note that the conversation generation program, conversation generation method, and conversation generation apparatus disclosed in the present application are not limited by the following embodiments.
Embodiment
[0014] FIG. 1 is a diagram showing an overview of a training system. The training system 1 includes a conversation generation apparatus 10 that conducts conversations as a virtual speaker and is responsible for training the target person P, and a target person terminal device 20 operated by the target person P. The conversation generation apparatus 10 and the target person terminal device 20 are connected via a network.
[0015] The conversation generation apparatus 10 is given information on a training scenario for defining how the conversation generation apparatus 10 behaves in the training for the target person P. The training scenario includes, for example, the background of the conversation, as well as the situation, mood, intention, and attitude of the conversation partner of the target person P. Then, the conversation generation apparatus 10 generates a statement according to the given training scenario and transmits the generated statement to the target person terminal device 20. Further, the conversation generation apparatus 10 generates a statement in consideration of the flow of the conversation for the statement of the target person P received from the target person terminal device 20 and transmits it to the target person terminal device 20. The conversation generation apparatus 10 thus acts as a pseudo-speaker for the target person P. The conversation generation apparatus 10 determines whether to end the conversation and ends the conversation according to the determination result.
[0016] The target person P is taught the purpose of the training and a training scenario similar to that given to the conversation generation apparatus 10. Then, the target person P uses the target person terminal device 20 to have a conversation with the conversation generation apparatus 10 that acts as a pseudo-speaker. The target person P receives training about the conversation through a role-play of having a conversation with the conversation generation apparatus 10 that acts as a pseudo-speaker as a conversation partner.
[0017] FIG. 2 is a block diagram of the conversation generation device. As shown in FIG. 2, the conversation generation device 10 includes a receiving unit 11, a data storage unit 12, a response output unit 13, a control unit 14, a first LLM 15, a prompt selection unit 16, a second LLM 17, and a database 18.
[0018] The data storage unit 12 stores the setting information 122 and the conversation history 121. The setting information 122 is information representing a training scenario. The setting information 122 includes information such as the background of the conversation, the situation, mood, intention, and attitude of the conversation partner of the target person P. The setting information 122 also includes information on the policy of the behavior of the conversation partner of the target person P and the information on the final state of the conversation. The final state of the conversation is set, for example, to either a state where the conversation ends with satisfaction or a state where a request to call a responsible person is made due to dissatisfaction.
[0019] This setting information 122 is an example of the situation setting information. For example, the situation setting includes information such as the conversation between the customer and the operator during claim handling, the handling of customer harassment, and the customer being angry. In this case, claim handling is an example of a "specific topic".
[0020] The conversation history 121 is the history of the conversation between the target person P and the conversation generation device 10 acting as a pseudo speaker. At the start of the conversation between the target person P and the pseudo speaker, the information is not registered, and as the conversation progresses, the conversation content is sequentially registered.
[0021] The database 18 stores a plurality of prompts 181. The prompt 181 is an instruction to the second LLM 17 including intention information indicating the intention of the next utterance of the conversation partner in the conversation in response to the utterance of the target person P. The intention information includes, for example, the emotion, action, request, etc. of the conversation partner in response to the utterance of the target person P. The prompt 181 is created in correspondence with the evaluation results obtained from various utterances. The evaluation of the utterance will be described later.
[0022] The receiving unit 11 receives the subject P's response to the simulated speaker's statement transmitted from the conversation generation device 10 from the subject terminal device 21. Here, the receiving unit 11 may receive text data representing the subject P's statement as a sentence, or it may receive audio data of the subject P's statement.
[0023] The receiving unit 11 registers the statements of the target person P at the end of the conversation in the conversation history 121 held by the data storage unit 12. This process of registering the conversation content in the conversation history 121 by the receiving unit 11 is an example of "processing to maintain a history of conversations with the target person." The receiving unit 11 also outputs the statements of the target person P to the control unit 14.
[0024] The first LLM15 is a large-scale language model that takes configuration information 122, conversation history 121, and subject P's statements as input, infers the evaluation value of subject P's statements, and outputs the inferred evaluation value as the evaluation result. The first LLM15 holds information on predetermined evaluation items. The first LLM15 determines whether the content of subject P's statements, which are input, contains content that satisfies each evaluation item, taking into account the input configuration information 122 and conversation history 121. The first LLM15 then assigns an evaluation value of 1 to evaluation items that contain that content, and an evaluation value of 0 to evaluation items that do not contain that content. The first LLM15 then outputs the combination of evaluation values for each evaluation item as the evaluation result.
[0025] Figure 3 shows an example of the evaluation results by the first LLM. For example, the first LLM15 has three evaluation items: "Does it include an apology?", "Does it show empathy?", and "Does it offer a follow-up suggestion?".
[0026] The first LLM 15 receives a prompt for evaluation value inference from the control unit 14, which includes setting information 122, conversation history 121, and statements made by subject P. If the statements made by subject P are audio data, the first LLM 15 performs speech recognition processing on the audio data and converts it into text data.
[0027] The first LLM15 receives input from subject P, such as, "I've been experiencing a problem with charging since yesterday, and I understand that it was repaired three months ago. I apologize for the inconvenience." In this case, the first LLM15 determines from the configuration information 122, conversation history 121, and subject P's statement that subject P's statement includes an apology and shows empathy, but does not offer any further suggestions. Then, as shown in evaluation result 101, the first LLM15 assigns a score of 1 to the evaluation item for whether an apology is included, a score of 1 to whether empathy is shown, and a score of 0 to whether any further suggestions are made. After that, the first LLM15 outputs "110" as the evaluation result, which is a combination of the evaluation values for each evaluation item.
[0028] Furthermore, the first LLM15 determines whether or not to terminate the conversation. For example, the first LLM15 has a predetermined upper limit on the number of responses. The first LLM15 then determines whether the number of responses from the simulated speaker has reached the upper limit. If the upper limit has been reached, the first LLM15 determines that the conversation is over. If the number of responses from the simulated speaker has not reached the upper limit, the first LLM15 determines whether the conversation situation has reached the final state included in the setting information 122. If the conversation situation has reached the final state, the first LLM15 determines that the conversation is over. If it determines that the conversation is over, it notifies the control unit 14 of the end of the conversation.
[0029] In this embodiment, the three evaluation items mentioned above were listed, but other types of evaluation items may be used. For example, other types of evaluation items such as "Is a firm attitude being maintained?" or "Is a distance being maintained from the other party?" can also be used. Furthermore, it is preferable to use different types of evaluation items depending on the scenario.
[0030] Furthermore, while the first LLM15 used binary values of 0 or 1 as evaluation values in this embodiment, it is also possible to use other values as evaluation values. For example, the first LLM15 may output the evaluation value for each evaluation item as text. In this case, the first LLM15 can output text such as "Did not apologize, not empathetic attitude, not concrete proposal" as the evaluation result.
[0031] Furthermore, for example, the first LLM15 can include intensity in the evaluation values for each item. The first LLM15 can represent evaluation values that include intensity for each evaluation item, for example, using values from 0 to 9. Specifically, if subject P's statement is polite, shows a strong apology, includes empathy, but does not include a suggestion, the first LLM15 can output an evaluation result of "950". In this case, the intensity can be inferred by the first LLM15 from subject P's statement, or the intensity information set for each keyword can be provided to the first LLM15, and the intensity can be calculated based on the settings according to the keyword.
[0032] Here, the first LLM15 is an example of a "speech content evaluation unit." In this embodiment, the first LLM15 was used as the speech content evaluation unit to infer the evaluation value of the subject P's speech using natural language processing, but the method of calculating the evaluation value by the speech content evaluation unit is not limited to this. For example, it is also possible to determine the evaluation value using a speech content evaluation unit that performs rule-based judgment processing, such as determining the evaluation value according to the keywords contained in the subject P's speech.
[0033] Furthermore, an example of the "first process" is the process of inputting setting information 122, conversation history 121, and the statements of subject P into the first LLM 15, inferring an evaluation value for the statements of subject P, and outputting the inferred evaluation value as the evaluation result. In other words, the statement content evaluation unit performs a first process that evaluates each of the pre-set evaluation items for the statements of subject P in a conversation about a specific topic for which a situation setting has been defined.
[0034] Returning to Figure 2, the explanation continues. The prompt selection unit 16 performs rule-based determination to select a prompt 181 from among multiple prompts 181 according to the evaluation result obtained by the first LLM 15. The selection of a prompt 181 by this prompt selection unit 16 branches the conversation scenario between the subject P and the virtual speaker. Hereafter, the rule-based prompt 181 determination process by the prompt selection unit 16 may be referred to as the "scenario branching determination process". The details of the prompt selection unit 16 are described below.
[0035] Figure 4 shows an example of a correspondence table that associates evaluation results with prompts. The prompt selection unit 16 holds, for example, the correspondence table 161 shown in Figure 4. The correspondence table 161 registers combinations of evaluation results, which are combinations of evaluation values for each evaluation item that are output by the first LLM 15, and the corresponding prompts 181. In Figure 4, different prompts 181 are represented as prompts #1 to #n. Each prompt 181 contains intention information that indicates the virtual speaker's intention in the next conversation. Intention information is information that determines the attitude the virtual speaker will take in the conversation, and includes, for example, the virtual speaker's current emotions during the conversation, the direction of the next action to be taken, and requests to the target person P. Each prompt 181 contains different intention information. Then, in the correspondence table 161, for a specific evaluation result, a prompt 181 containing intention information that the virtual speaker is expected to adopt based on the content of the statement corresponding to that specific evaluation result is associated and registered. This correspondence table 161 is an example of "information that associates combinations of evaluations for each evaluation item with intention information indicating the intentions of the person the subject is talking to in the next conversation."
[0036] The prompt selection unit 16 receives the evaluation result, which is a combination of evaluation values for each evaluation item, from the first LLM 15. Next, the prompt selection unit 16 refers to the correspondence table 161 and determines the prompt 181 that corresponds to the combination of evaluation values for each evaluation item included in the evaluation result. For example, if the evaluation result is "110", the prompt selection unit 16 refers to the correspondence table 161 and determines that prompt #k is the prompt 181 that corresponds to "110".
[0037] Next, the prompt selection unit 16 selects and retrieves the determined prompt 181 from among the multiple prompts 181 held by the database 18. For example, if the identified prompt 181 is prompt #k, the prompt selection unit 16 selects and retrieves prompt #k from among the multiple prompts 181 held by the database 18. After that, the prompt selection unit 16 outputs the selected prompt 181 to the control unit 14.
[0038] This prompt selection unit 16 is an example of a "selection unit." Furthermore, since each prompt 181 contains different intent information, selecting a prompt 181 is equivalent to selecting intent information. That is, the prompt selection unit 16 refers to a correspondence table 161 that associates combinations of evaluations for each evaluation item with intent information indicating the intent of the conversation partner of subject P in the next conversation. The prompt selection unit 16 then selects intent information corresponding to the combination of evaluations for each evaluation item obtained in the first processing by the first LLM 15.
[0039] Returning to Figure 2, the explanation continues. The second LLM17 is a large-scale language model that takes configuration information 122, conversation history 121, and prompt 181 as input, generates a response to the subject P's statements through inference, and outputs the generated response.
[0040] The second LLM 17 receives a response inference prompt from the control unit 14, in which the contents of the configuration information 122 and the conversation history 121 are added to the prompt 181. For example, the second LLM 17 receives a response inference prompt based on prompt 181, which has intent information consisting of feelings of frustration, actions to express frustration towards the user, and a request for a suggestion. The second LLM 17 then uses the configuration information 122, the conversation history 121, and the intent information to generate a response with content such as, "I was given the runaround last time, and it's the same content as last time, so I don't want to explain it again. What are you going to do about it?" The second LLM 17 then outputs the generated response to the response output unit 13.
[0041] Here, the second LLM 17 may generate the response as a text message and output it to the response output unit 13. Alternatively, the second LLM 17 can perform speech generation processing on the response to convert it into a speech message and output it to the response output unit 13. By converting the speech data of the subject P received by the first LLM 15 into text data, and the second LLM 17 converting the response into speech data, voice dialogue becomes possible between the subject P and the simulated speaker.
[0042] Here, the second LLM17 is an example of a "response generation unit." In this embodiment, the second LLM17 was used as the response generation unit to generate a response to the subject P's statement using natural language processing, but the method of generating a response by the response generation unit is not limited to this. For example, it is also possible to generate a response using rule-based judgment processing with a response generation unit configured to select a response from a plurality of pre-created response variations according to keywords included in the subject P's statement.
[0043] Furthermore, the response generation unit generates response content based on the intent information selected by the prompt selection unit 16, which is the selection unit, and the setting information 122, which is the situation setting information.
[0044] The response output unit 13 acquires the response to the statement made by the subject P, which is output by the second LLM 17. The response output unit 13 registers the acquired response at the end of the conversation in the conversation history 121 held by the data storage unit 12. This process of registering the conversation content in the conversation history 121 by the response output unit 13 is an example of "processing to maintain a history of conversations with the subject." The response output unit 13 also transmits the acquired response to the subject terminal device 30 as a statement made by the virtual speaker.
[0045] Here, the response output unit 13 can transmit text data and audio data of the simulated speaker's statements to the target terminal device 20. Furthermore, if the response output unit 13 acquires audio data from the second LLM 17, it can output it as audio and present it to the target P.
[0046] The control unit 14 receives input from the receiving unit 11, which is the subject P's statement. Next, the control unit 14 obtains the conversation history 121 and setting information 122 from the data storage unit 12. Then, the control unit 14 uses the conversation history 121, setting information 122, and the subject P's statement to generate an evaluation value inference prompt for the first LLM 15 to infer an evaluation value of the content of the subject P's statement. After that, the control unit 14 sends the generated evaluation value inference prompt to the first LLM 15.
[0047] Furthermore, the control unit 14 receives input of a prompt 181 corresponding to the evaluation result inferred by the first LLM 15 from the prompt selection unit 16. Next, the control unit 14 acquires the conversation history 121 and setting information 122 from the data storage unit 12. Then, the control unit 14 updates the acquired prompt 181 using the conversation history 121 and setting information 122 to generate a response inference prompt. After that, the control unit 14 feeds the generated response inference prompt to the second LLM 17.
[0048] Here, if the prompt 181 is configured to include an instruction to perform inference using the conversation history 121 and the configuration information 122, the control unit 14 may input the acquired prompt 181 directly to the second LLM 17 along with the conversation history 121 and the configuration information 122. Furthermore, although the phone embodiment was described as a configuration in which the conversation history 121 is updated by the receiving unit 11 and the response output unit 13, it is not limited to this. For example, the control unit 14 may register both the statements of the subject P and the statements of the virtual speaker in the conversation history 121, or the statements of the virtual speaker output from the second LLM 17 may be registered in the conversation history 121.
[0049] Furthermore, if the first LLM 15 determines that the conversation has ended, the control unit 14 receives notification of the end of the conversation from the first LLM 15. The control unit 14 then terminates the conversation between the subject P and the simulated speaker. This concludes the training for subject P.
[0050] In this embodiment, the first LLM 15 determined the end of the conversation, but the determination of the end of the conversation may be made by a unit other than the first LLM 15. For example, the prompt selection unit 16 may determine whether the upper limit of the number of responses has been reached, or whether the final state has been reached based on the evaluation value of the subject P's statements. In that case, the control unit 14 will terminate the conversation upon receiving notification of the end of the conversation from the prompt selection unit 16. Alternatively, the second LLM 17 may also determine the end of the conversation. In that case, the control unit 14 will terminate the conversation upon receiving notification of the end of the conversation from the second LLM 17.
[0051] In this way, the control unit 14 causes the first LLM 15, which is a speech content evaluation unit, to execute the first process. As the first process, the control unit 14 may input the speech of the subject P to the first LLM 15 and obtain evaluations for each evaluation item output from the first LLM 15 based on the content of the speech of the subject P. Alternatively, as the first process, if text data is input from the subject terminal device 20, the control unit 14 may acquire text data indicating the speech of the subject P, input the acquired text data to the first LLM 15, and output evaluations for each evaluation item. Furthermore, as the first process, the control unit 14 may perform the following process. For example, if audio data is input from the subject terminal device 20, the control unit 14 may acquire audio data of the speech of the subject P and input the acquired audio data to the first LLM 15. Then, the control unit 14 converts it into text data using speech recognition processing and outputs evaluations for each evaluation item based on the converted text data.
[0052] Furthermore, the control unit 14 causes the second LLM 17, which is a response generation unit, to perform the process of generating response content. As part of the process of generating response content, the control unit 14 may input the selected intent information and situation setting information to the second LLM 17 and obtain the response content output from the second LLM 17. Alternatively, as part of the process of generating response content, the control unit 14 may output text data indicating the response content, or it may perform a speech generation process on the text data indicating the response content to generate and output speech data. Furthermore, as part of the process of generating response content, the control unit 14 may generate response content based on the intent information included in the selected prompt 181, the setting information 122 which is situation setting information, and the conversation history 121.
[0053] Figure 5 is a diagram illustrating the overview of the processing performed by the conversation generation device in Example 1 of the training. Now, referring to Figure 5, the overall picture of the processing performed by the conversation generation device 10 in the training for subject P will be explained.
[0054] The receiving unit 11 acquires the statement from the target person P transmitted from the target person terminal device 20 (step S1).
[0055] The control unit 14 generates an evaluation value inference prompt from the conversation history 121, setting information 122, and the subject P's statements and inputs it to the first LLM 15. The first LLM 15 takes the conversation history 121, setting information 122, and the subject P's statements as input and performs language model processing to output an evaluation value of the content of the subject P's statements (step S2).
[0056] The prompt selection unit 16 receives an evaluation result input from the first LLM 15, which is a combination of evaluation values for the content of what the subject P said. Next, the prompt selection unit 16 determines the prompt 181 according to the evaluation result and executes a rule-based scenario branching determination process to select the prompt 181 from the database 18 according to the determination result (step S3).
[0057] The control unit 14 updates the prompt 181 selected by the prompt selection unit 16 using the conversation history 121 and setting information 122 to generate a response inference prompt and inputs it to the second LLM 17. The second LLM 17 takes the conversation history 121, setting information 122, and intent information contained in the prompt 181 as input and performs language model processing to output a response to the subject P's statement (step S4).
[0058] If the conversation continues, the operation of the conversation generation device 10 returns to step S1, and steps S1 to S4 are repeated.
[0059] Figure 6 is a flowchart of the training process for a subject using the conversation generation device according to Example 1. Next, referring to Figure 6, the flow of the training process for subject P using the conversation generation device 10 according to Example 1 will be explained.
[0060] Subject P transmits their statements to the conversation generation device 10 using the subject terminal device 20. The receiving unit 11 acquires the statements from subject P transmitted from the subject terminal device 20 (step S11).
[0061] The control unit 14 generates an evaluation value inference prompt from the conversation history 121, setting information 122, and the statements of the subject P, and inputs it to the first LLM 15 (step S12).
[0062] The first LLM15 takes the conversation history 121, setting information 122, and the statements of subject P as input to perform inference, and outputs an evaluation result by combining the evaluation values of the content of subject P's statements (step S13).
[0063] The prompt selection unit 16 receives an evaluation result from the first LLM 15, which is a combination of evaluation values for the content of what subject P said. Next, the prompt selection unit 16 determines a prompt 181 corresponding to the evaluation result and selects a prompt 181 from the database 18 according to the determination result (step S14).
[0064] The control unit 14 receives the input of the selected prompt 181 from the prompt selection unit 16. Next, the control unit 14 updates the selected prompt 181 using the conversation history 121 and setting information 122 to generate a response inference prompt. Then, the control unit 14 feeds the response inference prompt to the second LLM 17 (step S15).
[0065] The second LLM 17 takes the conversation history 121, setting information 122, and intent information contained in the prompt 181 as input and performs inference, outputting a response to the subject P's statement. The response output unit 13 sends the response output from the second LLM 17 to the subject terminal device 20 as the statement of the pseudo-speaker, presenting the response to the subject P (step S16).
[0066] The first LLM 15 determines whether to continue the conversation (step S17). If the conversation is to continue (step S17: affirmative), the training process returns to step S11. Conversely, if the conversation is to end (step S17: negative), the first LLM 15 notifies the control unit 14 that the conversation has ended. Upon receiving notification of the end of the conversation, the control unit 14 terminates the conversation between subject P and the simulated speaker. This completes subject P's training.
[0067] As described above, the conversation generation device 10 according to this embodiment calculates an evaluation value of the content of the subject P's statements, and selects a prompt 181 containing intent information that the conversation partner is expected to adopt according to the calculated evaluation value, thereby branching the scenario. In this way, by preparing prompts containing intent information corresponding to the flow of the conversation and the attitude of subject P in advance, and selecting prompts 181 according to the content of subject P's statements, it is possible to realize scenario branching according to the flow of the conversation and the attitude of subject P. Therefore, by returning an appropriate response according to the conversation situation, natural conversation can be achieved. In other words, it is possible to realize role-playing that assumes complex scenarios including changes in the intent of the simulated speaker, and it is possible to provide effective conversation training. [Examples]
[0068] Next, we will describe Example 2. The conversation generation device 10 according to this example is also represented by the block diagram in Figure 2. The conversation generation device 10 according to this example reflects the evaluation based on the audio data of the subject P's statements in the evaluation value of the subject P's statements. In the following description, the operation of each part, as in Example 1, may be omitted.
[0069] The receiving unit 11 receives audio data of the subject P's statements from the subject terminal device 20. The receiving unit 11 then outputs the audio data of the subject P's statements to the control unit 14.
[0070] The control unit 14 generates an evaluation value inference prompt from the conversation history 121, setting information 122, and the audio data of the subject P's statements, and inputs it to the first LLM 15.
[0071] The first LLM15 performs speech analysis on the audio data of subject P's speech to obtain information about the characteristics of the conversation. For example, the first LLM15 obtains information about the characteristics of the conversation, such as subject P's tone of voice, the time interval between responses, and the speed of the conversation.
[0072] Furthermore, the first LLM15 processes the audio data using speech recognition and other methods to convert it into text data and obtain the content of the speech. Then, the first LLM15 infers predetermined evaluation values for each item regarding the speech of subject P from the content of the speech and the characteristics of the conversation. Finally, the first LLM15 outputs an evaluation result that combines the evaluation values for each item.
[0073] Thus, as a first process, the control unit 14 may input the audio data of the subject P's statements to the first LLM 15, which is a statement content evaluation unit, to extract the characteristics of the subject P's conversation from the audio data, and output an evaluation for each evaluation item based on the characteristics of the subject P's conversation and the content of the statements.
[0074] As described above, the conversation generation device 10 according to this embodiment calculates an evaluation value of the subject P's statements, taking into account the content of the statements and the characteristics of the conversation, and branches the scenario accordingly. By using the characteristics of the conversation in evaluating the statements of the subject P in this way, it is possible to appropriately select intention information that is more in line with the flow of the conversation and the attitude of the subject P, thereby realizing more appropriate scenario branching that is in line with the flow of the conversation and the attitude of the subject P. Therefore, it is possible to realize more natural conversations and provide more effective conversation training through role-playing that assumes complex scenarios including changes in the intentions of the simulated speaker.
[0075] (Hardware configuration) Figure 7 is a hardware configuration diagram of the conversation generation device. Next, an example of a hardware configuration for realizing each function of the conversation generation device 10 will be described with reference to Figure 7.
[0076] As shown in Figure 7, the conversation generation device 10 includes, for example, a CPU (Central Processing Unit) 91, memory 92, a hard disk 93, and a network interface 94. The CPU 91 is connected to the memory 92, hard disk 93, and network interface 94 via a bus.
[0077] The network interface 94 is an interface for communication between the conversation generation device 10 and an external device. For example, the network interface 94 relays communication between the target terminal device 20 and the CPU 91.
[0078] The hard disk 93 is an auxiliary storage device. The hard disk 93 implements the functions of the data storage unit 12 and the database 18 as illustrated in Figure 2. The hard disk 93 may also store the first LLM 15 and the second LLM 17. Furthermore, the hard disk 93 stores various programs, including programs for implementing the functions of the receiving unit 11, data storage unit 12, response output unit 13, control unit 14, and prompt selection unit 16 as illustrated in Figure 2.
[0079] Memory 92 is the main memory. Memory 92 can be, for example, DRAM (Dynamic Random Access Memory).
[0080] The CPU 91 reads various programs from the hard disk 93, loads them into memory 92, and executes them. This allows the CPU 91 to implement the functions of the receiving unit 11, data storage unit 12, response output unit 13, control unit 14, and prompt selection unit 16, as illustrated in Figure 2. [Explanation of Symbols]
[0081] 1. Training System 10 Conversation generation device 11 Receiving unit 12 Data storage unit 13 Response Output Section 14 Control Unit 15 1st LLM 16. Prompt Selection Section 17 2nd LLM 18 Databases 20 Target user terminal device 121 Conversation History 122 Configuration Information 181 Prompt
Claims
1. The first process involves performing an evaluation of the statements made by the participants in a conversation about a specific topic with a defined context, according to pre-defined evaluation criteria. By referring to information that associates the combination of evaluations for each of the evaluation items with intention information indicating the intention of the person the subject is talking to in the next conversation, the intention information corresponding to the combination of evaluations for each of the evaluation items obtained in the first process is selected. The response content is generated based on the selected intent information and the status setting information. A conversation generation program characterized by having a computer perform the processing.
2. The first process includes inputting the subject's statements into the statement content evaluation unit and obtaining evaluations for each evaluation item output from the statement content evaluation unit based on the content of the subject's statements. The process for generating the response content includes inputting the selected intent information and the status setting information into the response generation unit and obtaining the response content output from the response generation unit. The conversation generation program according to feature 1.
3. The first process includes acquiring text data representing the subject's statements, inputting the acquired text data into the statement content evaluation unit, and outputting an evaluation for each evaluation item. The process for generating the response content includes a process for causing the response generation unit to output text data indicating the response content. The conversation generation program according to feature 2.
4. The first process includes acquiring audio data of the subject's statements, inputting the acquired audio data into the statement content evaluation unit, converting it into text data through speech recognition processing, and outputting an evaluation for each evaluation item based on the converted text data. The process for generating the response content includes a process in which the response generation unit performs a speech generation process on the text data representing the response content to generate and output speech data. The conversation generation program according to feature 2.
5. The conversation generation program according to claim 2, characterized in that the first process includes inputting audio data of the subject's statements into the statement content evaluation unit, extracting characteristics of the subject's conversation from the audio data, and outputting an evaluation for each evaluation item based on the characteristics of the subject's conversation and the content of the statements.
6. The computer is further instructed to perform a process to retain a history of conversations with the aforementioned subject. The first process includes a process of performing an evaluation for each evaluation item based on the subject's statements and history, The process for generating the response content includes a process for generating the response content based on the selected intent information, the status setting information, and the history. The conversation generation program according to feature 1.
7. The conversation generation device, The first process involves performing an evaluation of the statements made by the participants in a conversation about a specific topic with a defined context, according to pre-defined evaluation criteria. By referring to information that associates the combination of evaluations for each of the evaluation items with intention information indicating the intention of the person the subject is talking to in the next conversation, the intention information corresponding to the combination of evaluations for each of the evaluation items obtained in the first process is selected. The response content is generated based on the selected intent information and the status setting information. A conversation generation method characterized by performing a process.
8. A statement content evaluation unit performs a first process that evaluates the statements made by the subject of a conversation on a specific topic with a defined context, according to pre-defined evaluation items. A selection unit selects the intent information corresponding to the combination of evaluations for each evaluation item obtained in the first processing by the utterance content evaluation unit, by referring to information relating the combination of evaluations for each evaluation item and intent information indicating the intent of the person the subject is talking to in the next conversation. A response generation unit generates response content based on the intent information and the status setting information selected by the selection unit. A conversation generation device characterized by being equipped with the following features.