Voice evaluation device and voice evaluation method

The voice evaluation device uses a generative AI model to generate unbiased second evaluation information, addressing bias and reducing costs in voice evaluation systems.

WO2026028299A1PCT designated stage Publication Date: 2026-02-05NTT DOCOMO INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/027185
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Existing voice evaluation systems are biased towards the evaluator's viewpoint, leading to inconsistent quality in evaluation results, and increasing personnel costs with multiple evaluators.

Method used

A voice evaluation device using a generative AI model generated by machine learning processes voice information and first evaluation information to generate unbiased second evaluation information, reducing the need for multiple human evaluators.

Benefits of technology

Ensures high-quality evaluation results while minimizing personnel costs by providing a more comprehensive and unbiased assessment of voice quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024027185_05022026_PF_FP_ABST
    Figure JP2024027185_05022026_PF_FP_ABST
Patent Text Reader

Abstract

A voice evaluation device 10 comprises: an acquisition unit 20 that acquires voice information on the voice of a subject to be evaluated and first evaluation information D1 indicating a first evaluation of the voice and a reason for the first evaluation; and a processing unit 30 that executes processing on the voice information and the first evaluation information D1 to obtain second evaluation information D5 indicating a second evaluation of the voice based on the voice information and the first evaluation information D1 by a generative AI model L generated by machine learning.
Need to check novelty before this filing date? Find Prior Art

Description

Audio evaluation device and audio evaluation method

[0001] The present invention relates to a voice evaluation device and a voice evaluation method.

[0002] Conventionally, voice evaluation devices that evaluate the voice of an evaluation target person are known. Patent Document 1 describes a voice evaluation device that evaluates an operator's telephone response. In this voice evaluation device, a voice file indicating the operator's telephone response is judged based on pre-registered evaluation conditions, and the result is generated as a voice evaluation result.

[0003] JP 2014-86942 A

[0004] However, in the above-mentioned voice evaluation device, the evaluation conditions for evaluating the voice are registered in advance by the user of the voice evaluation device, and therefore the viewpoint of evaluation of the voice of the person to be evaluated is biased, which may result in the quality of the evaluation results of the voice of the person to be evaluated not being guaranteed.

[0005] On the other hand, when multiple evaluators evaluate the speech of a person being evaluated, the more evaluators there are, the more viewpoints the speech can be evaluated from, and the less bias there is in the evaluation viewpoints. This ensures the quality of the evaluation results of the person being evaluated, but it also increases the personnel costs.

[0006] One embodiment of the present invention has been made in consideration of the above, and aims to provide a voice evaluation device and a voice evaluation method that can reduce human costs while ensuring the quality of the evaluation results of the voice of the person being evaluated.

[0007] In order to achieve the above-mentioned object, a voice evaluation device according to one embodiment of the present invention comprises an acquisition unit that acquires voice information relating to the voice of a person to be evaluated, as well as first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation, and a processing unit that executes processing on the voice information and the first evaluation information to obtain second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information using a generative AI model generated by machine learning.

[0008] A voice evaluation device according to one embodiment of the present invention performs processing on the voice information and the first evaluation information to obtain, using a generative AI model generated by machine learning, second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation. According to this configuration, the generative AI model generates the second evaluation information based at least on the reason for the first evaluation in the first evaluation information. This reduces bias in the reason for the second evaluation in the second evaluation information compared to bias in the reason for the first evaluation in the first evaluation information, thereby ensuring the quality of the second evaluation information. As a result, the quality of the evaluation results of the voice of the subject of evaluation can be ensured. Furthermore, bias in the reason for the second evaluation in the second evaluation information can be reduced without increasing the number of pieces of first evaluation information. This ensures the quality of the second evaluation information while reducing the personnel costs required to increase the number of pieces of first evaluation information. As a result, it is possible to reduce personnel costs while ensuring the quality of the evaluation results of the voice of the subject of evaluation.

[0009] Incidentally, one embodiment of the present invention can be described not only as an apparatus invention as described above, but also as a method invention as described below. These are essentially the same inventions, just in different categories, and have similar functions and effects.

[0010] That is, a voice evaluation method according to one embodiment of the present invention comprises an acquisition step of acquiring voice information relating to the voice of a person to be evaluated, as well as first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation, and a processing step of executing processing on the voice information and the first evaluation information to obtain second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information using a generative AI model generated by machine learning.

[0011] According to one embodiment of the present invention, it is possible to reduce personnel costs while ensuring the quality of the evaluation results of the voice of the person being evaluated.

[0012] 1A, 1B, 1C, and 1D are diagrams for explaining an overview of a voice evaluation device according to an embodiment; FIG. 1A is a block diagram showing an information processing system according to an embodiment; FIG. 1B is a diagram showing an example of first evaluation information; FIG. 1C is a diagram showing an example of a prompt for causing a generation AI model to generate trend information; FIG. 1D is a diagram showing an example of trend information; FIG. 1E is a diagram showing an example of a prompt for causing a generation AI model to generate third evaluation information; FIG. 1F is a diagram showing an example of the third evaluation information; FIG. 1F is a diagram showing an example of a prompt for causing a generation AI model to generate second evaluation information; FIG. 1F is a diagram showing an example of second evaluation information; FIG. 1F is a flowchart showing an example of processing of a voice evaluation method according to an embodiment; FIG. 1F is a diagram showing an example of a prompt for generating second evaluation information in a modified example; FIG. 1F is a diagram showing another example of a prompt for generating second evaluation information in a modified example; FIG. 1F is a diagram showing an example of second evaluation information in a modified example;

[0013] Hereinafter, an embodiment of a voice evaluation device and a voice evaluation method according to the present invention will be described in detail with reference to the drawings. In the description of the drawings, the same elements are given the same reference numerals and duplicated explanations will be omitted.

[0014] Conventionally, the voice of a person being evaluated has been evaluated. For example, in a company that provides products or services to customers, operators working at a call center respond to customer inquiries over the phone. In such a company, in order to improve customer reputations, evaluations of recorded voices of the operator's customer service (e.g., a conversation between the operator and a customer) are fed back to the operator. As an example, the recorded voices of the operator's customer service are evaluated by multiple evaluators (e.g., the operator's supervisor and other managers). In this case, the multiple evaluators have different values ​​and therefore evaluate the voices for different reasons (e.g., perspectives or grounds). Therefore, as the number of evaluators increases, the number of reasons for evaluating the voices increases, and bias in the reasons for evaluating the voices is reduced. Then, by feeding back evaluation results in which bias in the reasons for evaluating the voices is reduced to the operators, the quality of the operator's customer service can be more reliably improved. However, there has been a problem in that the human resources costs required for evaluating the voices increase as the number of evaluators is increased in order to reduce bias in the reasons for evaluating the voices. The human cost includes, for example, the number of people involved in a certain task and the time each person spends on the task. In this embodiment, the reason for evaluation is, for example, the basis for evaluation or the viewpoint of evaluation.

[0015] In order to solve the above problem, the voice evaluation device according to this embodiment suppresses bias in the reasons for evaluating voice without increasing the human cost spent on evaluating voice. Specifically, the voice evaluation device causes a generative AI model to evaluate the voice of a subject for evaluation for reasons different from those of the multiple evaluators based on multiple evaluation results from multiple evaluators. This allows the generative AI model to evaluate the voice as, for example, a different evaluator with a different perspective from the multiple evaluators. As a result, the voice evaluation device can increase the number of actual evaluators without increasing human cost, and suppress bias in the reasons for evaluating voice. In other words, even when the number of evaluators is small, the voice evaluation device can suppress an increase in bias in the reasons for evaluating voice by evaluating voice using a generative AI model.

[0016] The processing in the voice evaluation device will be described. FIGS. 1A, 1B, 1C, and 1D are diagrams for explaining an overview of the processing in the voice evaluation device. First, in the example shown in FIG. 1A, the voice of a subject (e.g., an operator) is evaluated by multiple evaluators E (e.g., the operator's superiors and other managers), and first evaluation information D1 is generated in advance for each evaluator E. The first evaluation information D1 is information indicating an evaluation of the subject's voice. In the example shown in FIG. 1B, the voice evaluation device generates trend information D2 indicating a trend in evaluations of the subject's voice based on the multiple pieces of first evaluation information D1. The trend in evaluation indicates the magnitude of bias in the reasons for evaluation of the multiple pieces of first evaluation information D1 (e.g., a result of evaluating the magnitude of bias in the reasons for evaluation on a three-point scale) and the trend in bias in the reasons for evaluation of the multiple pieces of first evaluation information D1 (e.g., a sentence explaining specifically how the reasons for evaluation are biased).

[0017] 1(c), the voice evaluation device generates a prompt D3 to be input to the generation AI model L based on the plurality of pieces of first evaluation information D1 and the trend information D2. The voice evaluation device inputs the generated prompt D3 to the generation AI model L to obtain third evaluation information D4 indicating an evaluation of the voice of the person being evaluated. The third evaluation information D4 is information indicating an evaluation of the voice of the person being evaluated. By generating the third evaluation information D4 based on the trend information D2 indicating a tendency of bias in the reasons for the evaluation of the plurality of pieces of first evaluation information D1, the evaluation in the third evaluation information D4 can be an evaluation having a reason different from the reasons for the evaluation of the plurality of pieces of first evaluation information D1.

[0018] 1D, second evaluation information D5 is generated based on the plurality of pieces of first evaluation information D1 and third evaluation information D4. As a result, the evaluation reasons of the second evaluation information D5 include both the evaluation reasons of the plurality of pieces of first evaluation information D1 and the evaluation reasons of the third evaluation information D4. As a result, the second evaluation information D5 includes more reasons, and therefore, the bias of the evaluation reasons in the second evaluation information D5 is suppressed.

[0019] The generative AI model L is a model that can generate content in response to input of a prompt (details of which will be described later) including input information, according to any one or a combination of the instructions, context, question, and output format indicated by the prompt, and return the content as response information. The generative AI model L generates response information targeted at the input information. The generative AI model L may be, for example, an interactive AI model that includes a large-scale language model (LLM) and a user interface (UI) for interacting with the user, enabling text chat or voice chat with the user. Examples of such generative AI models include ChatGPT, GPT (registered trademark)-3.5, GPT-4V, PaLM2, etc. Furthermore, although the above describes an example of a large-scale language model, other AI models may also be used.

[0020] A prompt is information indicating instructions or questions entered by a user in an interactive system, such as a dialogue with a generative AI model or a command line interface (CLI). The prompt uses text to express, for example, the command to be executed by the interactive AI model, the task to be executed by the interactive AI model, the background / context to be considered by the interactive AI model (e.g., role, condition), the question to be answered by the interactive AI model, and the output format of the response information from the interactive AI model. The prompt may also include input information that is the target of the command / task to be executed by the interactive AI. Examples of such input information include data files with file names that include a predetermined extension, such as text data, image data, application-related data, audio data, video data, and still image data. Application-related data is data such as document data, table data, and graph data that can be processed by a default application program.

[0021] As an example, the prompt includes an instruction for the generated AI model L. The instruction includes at least one of an explanation regarding the input information to be input to the generated AI model L and an explanation regarding the response information to be output by the generated AI model L.

[0022] Next, the functions of the speech evaluation device 10 according to this embodiment will be described. Fig. 2 is a block diagram showing an information processing system according to this embodiment. As shown in Fig. 2, the information processing system 1 includes a terminal A, a speech evaluation device 10, and a server device 100.

[0023] Terminal A is a device used by a user who wishes to obtain various types of content (text, audio, images, videos, etc.) based on input information using interactive AI. Terminal A is, for example, a personal computer, smartphone, tablet terminal, feature phone, server device, game console, etc. Note that although only one terminal A is illustrated in FIG. 2 , the information processing system 1 may include any number of terminals A, two or more.

[0024] The server device 100 is a device that enables the provision of content using the generated AI model L. In this embodiment, the server device 100 is capable of providing a content provision function using the generated AI model L. The generated AI model L may be stored within the server device 100, or may be stored in another device connected to the server device 100 via a network, and configured to enable information exchange with the voice evaluation device 10 via the server device 100.

[0025] The speech evaluation device 10 according to this embodiment is, for example, a Retrieval-Augmented Generation (RAG) system. The knowledge database searched by the RAG system is implemented, for example, within the speech evaluation device 10 (not shown). The speech evaluation device 10 includes an acquisition unit 20 and a processing unit 30. The functions of each functional unit of the speech evaluation device 10 will be described in detail below.

[0026] The acquisition unit 20 acquires voice information related to the voice of the person to be evaluated. Specifically, the acquisition unit 20 acquires the voice information from terminal A. In this embodiment, the voice information is voice data including human speech. For example, the voice information is data obtained by recording a conversation between an operator and a customer. Note that the voice information may also be voice information including other voices other than those described above.

[0027] The acquisition unit 20 acquires first evaluation information D1 indicating a first evaluation of the voice of the person to be evaluated (e.g., an operator) and the reason for the first evaluation (see FIG. 1A). Specifically, multiple evaluators (e.g., the operator's superiors and other managers) each create multiple pieces of first evaluation information D1 and input them to terminal A. The acquisition unit 20 acquires the multiple pieces of first evaluation information D1 from terminal A. In this embodiment, the first evaluation information D1 is information indicating a first evaluation of the voice of the person to be evaluated and the reason for the first evaluation for each item. The multiple items are set in advance by a user or the like according to the type of voice of the person to be evaluated. For example, the type of voice is a recording of a conversation between a call center employee and a customer. The multiple items are items related to the satisfaction level of customers who called the call center. The first evaluation is a first evaluation result indicating an evaluation of the voice of the person to be evaluated using one of multiple stages. The multiple levels indicate the level of quality of the evaluation of the subject's voice (for example, "A," "B," and "C" as shown in FIG. 3, where "A" indicates the best evaluation, "B" indicates the second best evaluation, and "C" indicates the worst evaluation). The reason for the first evaluation is, for example, a sentence explaining the reason.

[0028] FIG. 3 is a diagram illustrating an example of first evaluation information. In the example illustrated in FIG. 3, the first evaluation information includes information indicating a first evaluation in which the speech of the person being evaluated (the operator) is evaluated on a three-point scale ("A," "B," and "C") for each item. The first evaluation information also includes information indicating the reason for the first evaluation, which is the evaluator's comments for each item. The items are "opening," "confirmation of the purpose," "persuasiveness" (ability to explain), "speaking time and speed," "proposal," "voice expression," "acknowledgment / thanks," "language choice," and "closing." A good evaluation of "opening" indicates that the operator's speech when starting a conversation with a customer is appropriate. A good evaluation of "confirmation of the purpose" indicates that the operator repeats the customer's purpose and asks questions to clarify if the purpose is unclear. A good evaluation of "persuasiveness" (ability to explain) indicates that the operator explains concisely in words that the customer can understand. A good evaluation of "speaking time and speed" indicates that the operator speaks at a time and speed that is tailored to the customer. A good evaluation of "suggestion" indicates that the operator makes various suggestions to please the customer. A good evaluation of "voice expression" indicates that the operator speaks with a voice expression appropriate to the situation. A good evaluation of "backchannels and acknowledgments" indicates that both the way the operator responds to what the customer says and the way the operator thanks the customer for what they say are appropriate. A good evaluation of "language" indicates that the operator speaks appropriately, using business terms and honorific language. A good evaluation of "closing" indicates that the operator's speech at the end of the conversation with the customer is appropriate. In the following explanation, "confirming the purpose of the conversation," "ability to explain," "speaking time and speed," "voice expression," and "language" are used as multiple items.

[0029] The processing unit 30 processes the voice information and the first evaluation information D1 to obtain second evaluation information D5 based on the voice information and the first evaluation information D1 using the generation AI model L (see Figures 1(b), 1(c), and 1(d)). The second evaluation information D5 is information indicating a second evaluation of the voice of the person being evaluated and the reason for the second evaluation (details will be described later). In this embodiment, the processing unit 30 inputs input information based on the voice information and the first evaluation information D1 to the generation AI model L to obtain the second evaluation information D5. For example, the processing unit 30 generates a prompt D7 (described later) to be input to the generation AI model based on the voice information and the first evaluation information D1. The processing unit 30 inputs the generated prompt D7 to the generation AI model L to obtain the second evaluation information D5.

[0030] The second evaluation information D5 is information indicating a second evaluation of the voice of the person being evaluated and the reason for the second evaluation for each item. The multiple items are set in advance by a user or the like according to the type of voice of the person being evaluated. For example, the multiple items are items related to the satisfaction of customers who called the call center. The multiple items related to the second evaluation are usually the same as the multiple items related to the first evaluation. The second evaluation is a second evaluation result indicating an evaluation of the voice of the person being evaluated using one of multiple stages. The stage related to the second evaluation may be the same as the stage related to the first evaluation. The reason for the second evaluation is, for example, a sentence explaining the reason. The second evaluation information D5 may be in the same format as the first evaluation information D1.

[0031] Specifically, first, the processing unit 30 generates trend information D2 based on the voice information and the first evaluation information D1 (see FIG. 1B). The processing unit 30 executes processing on the voice information and the first evaluation information D1 to obtain the trend information D2 based on the voice information and the first evaluation information D1 using the generation AI model L. The trend information D2 is information indicating the trend of the reasons for the first evaluations of one or more pieces of first evaluation information D1. For example, the trend information D2 is information indicating the magnitude of bias in the reasons for the first evaluations of one or more pieces of first evaluation information D1, and a sentence indicating which part of the voice information was the basis for each first evaluation, or a sentence describing the state of the person being evaluated in part or all of the voice information.

[0032] The bias in the reasons for the first evaluations in the plurality of pieces of first evaluation information D1 will be described. For example, if the plurality of pieces of first evaluation information D1 are evaluated from a specific biased perspective when they should have been evaluated from various perspectives, the bias in the reasons for the first evaluations is deemed to be large. Also, if the plurality of pieces of first evaluation information D1 are evaluated from a specific biased portion when they should have been evaluated from a wide range of audio, the bias in the reasons for the first evaluations is deemed to be large. The same applies to the bias in the reasons for the second evaluations in the plurality of pieces of second evaluation information D5.

[0033] For example, the processing unit 30 generates text information indicating the content of the speech of the person to be evaluated based on the speech information. More specifically, the acquisition unit 20 generates the text information by converting the speech indicated by the speech information into text. The process of converting speech into text is realized by various known methods, such as known speech recognition methods. Note that the processing unit 30 does not need to generate text information based on the speech information. In this case, the speech information is text data indicating the speech of the person to be evaluated, and is treated as text information in the following processing.

[0034] The processing unit 30 then generates trend information D2 based on the plurality of pieces of first evaluation information and text information. More specifically, it generates a trend evaluation prompt D6, which is a prompt for obtaining the trend information D2 based on the plurality of pieces of first evaluation information and text information using the generation AI model L. The trend information D2 is generated collectively for the plurality of pieces of first evaluation information, but may also be generated for each piece of first evaluation information. As an example, the processing unit 30 generates input information based on the plurality of pieces of first evaluation information and text information and includes the input information in a prompt (hereinafter referred to as a first general prompt) stored in the speech evaluation device 10, thereby generating the trend evaluation prompt D6.

[0035] 4 is a diagram showing an example of the trend evaluation prompt D6. In the example shown in FIG. 4, the first general prompt is a command statement including command statements D61, D62, and D63. The processing unit 30 inputs input information including a plurality of pieces of first evaluation information D1 and text information into input fields in the first general prompt, and generates command statements D64 and D65, thereby generating the trend evaluation prompt D6. In the first general prompt, character strings indicating that these are the fields to input the first evaluation information D1 and the text information are provided after the command statements D62 and D63, respectively.

[0036] The instruction D61 of the tendency evaluation prompt D6 includes information to generate tendency information D2 based on the text information and the first evaluation information D1. In the example shown in FIG. 4 , the information is the following: "#Instruction #The interaction history is a conversation between a staff member and a customer. Each line includes the line number, speaker ID, text of the speech recognition result, emotion recognition result, and speaking rate information. #The evaluator's evaluation result is the evaluator's evaluation result of the staff member's interaction. #Please refer to the interaction history and the evaluator's evaluation result, and analyze the magnitude and trend of bias in the reasons for the evaluation in the #evaluator's evaluation result. #The magnitude of bias in the reasons in the output format indicates the degree to which the points that should be the basis for the evaluation result for each item are covered in the #evaluator's evaluation result."

[0037] The above section includes information describing the text information. The information is the following: "# Response History is a conversation between a staff member and a customer. Each line includes a line number, a speaker ID, a text of the speech recognition result, an emotion recognition result, and speech rate information." The "staff member" is, for example, an operator. The line number is a number assigned to each line of # Response History. The speaker ID is an identification ID assigned to each speaker. The emotion recognition result is the result of recognizing the speaker's emotion for each utterance. The speech rate information is information indicating the speech rate for each utterance. The line number, speaker ID, text of the speech recognition result, the emotion recognition result, and speech rate information are generated from speech information using known techniques. The text information may include the above information. Furthermore, the tendency information D2 generated by the generation AI model L may be generated taking into account the emotion and speech rate associated with the speech.

[0038] The above section includes information explaining the first evaluation information D1. This information is the section "#Evaluation result by evaluator is the evaluation result by the evaluator of the staff's response." The evaluator may be, for example, the operator's superior or other manager.

[0039] The above portion includes information to generate trend information D2 based on the input information. Specifically, the instruction sentence D61 includes information to generate trend information D2 indicating the trend of evaluations in the plurality of pieces of first evaluation information D1 based on the plurality of pieces of first evaluation information D1 and text information. The information is the portion that reads, "Please refer to the # interaction history and the # evaluation results by the evaluator, and analyze the magnitude and trend of bias in the reasons for the evaluations in the # evaluation results by the evaluator."

[0040] The above section includes information explaining the bias in the reasons for the first evaluation indicated by the trend information D2. This information is the section that reads, "#The magnitude of bias in the reasons in the output format indicates the degree to which the points that should be the basis for the evaluation results for each item are covered in the evaluation results by the #evaluator."

[0041] The instruction sentence D62 of the tendency assessment prompt D6 includes content explaining each of the multiple items contained in the first assessment information D1. In the example shown in Fig. 4, the information is the following: "#Check items 1. Confirmation of purpose: The person repeats the customer's purpose. If it is not clear, the person asks questions to clarify. 2. Ability to explain: The person explains concisely in words that the customer can understand. 3. Pause and speed of speech: The person speaks at a pace and speed that suits the customer. 4. Voice expression: The person speaks with a voice expression appropriate for the situation. 5. Language: The person speaks using business terms and honorific language."

[0042] The instruction D63 of the tendency evaluation prompt D6 includes content indicating the format of the tendency information D2. For example, the instruction D63 may include content to the effect that the tendency information D2 includes, for each item, the magnitude of bias in the reasons for evaluation of the voice of the person being evaluated and a sentence explaining the tendency of the reasons for the evaluation. A "large" bias indicates that the reasons for evaluation are biased to a large extent, a "medium" bias indicates that the reasons for evaluation are biased to the second largest extent, and a "small" bias indicates that the reasons for evaluation are biased to a small extent. In the example shown in FIG. 4, the information is the following: "#Output format 1. Confirmation of purpose: [large / medium / small bias] Tendency: [describe the tendency of the reason for the evaluation] 2. Explanation ability: [large / medium / small bias] Tendency: [describe the tendency of the reason for the evaluation] 3. Speech timing and speed: [large / medium / small bias] Tendency: [describe the tendency of the reason for the evaluation] 4. Voice expression: [large / medium / small bias] Tendency: [describe the tendency of the reason for the evaluation] 5. Language: [large / medium / small bias] Tendency: [describe the tendency of the reason for the evaluation]."

[0043] The instruction sentence D64 of the tendency evaluation prompt D6 includes content indicating a plurality of pieces of first evaluation information D1. The instruction sentence D64 includes content indicating a plurality of items, an evaluation for each item, and a comment for each item, for each evaluator. In the example shown in FIG. 4 , the information is the following: "Evaluation results by #evaluators <Evaluation result of first evaluator> 1. Confirmation of the purpose: A Comment: The subject was able to repeat the purpose. 2. Explanation ability: B Comment: The sentence was long and redundant. 3. Speech timing and speed: A Comment: No particular problems. 4. Voice expression: A Comment: The subject spoke with an appropriate voice expression. 5. Language: B Comment: The subject used the expression "Was it okay?" <Evaluation result of second evaluator> ... <Evaluation result of third evaluator> ..."

[0044] The instruction D65 of the tendency evaluation prompt D6 includes the content of the text information, which is the "# response history..." portion.

[0045] The processing unit 30 inputs the generated tendency evaluation prompt D6 into the generation AI model L, thereby obtaining tendency information D2 generated by the generation AI model L. The tendency information D2 is information indicating the tendency of the reasons for the first evaluation of one or more pieces of first evaluation information D1. For example, the tendency information D2 includes, for each item, the magnitude of bias in the reasons for evaluation of the speech of the person being evaluated based on multiple pieces of first evaluation information and text information, and information indicating the tendency of the reasons for the evaluation. FIG. 5 is a diagram showing an example of the tendency information D2. In the example shown in FIG. 5, the tendency information D2 indicates, for each item, evaluations based on multiple pieces of first evaluation information and text information, and the tendency of the reasons for the evaluation. The information includes the following: 1. Confirmation of the purpose: Small tendency: Evaluation is carried out based on the repetition of the requirements in line 5, "..." 2. Explanation ability: Medium tendency: Evaluation is based on the length of the explanation and evaluation based on rephrasing 3. Speaking pause and speed: Medium tendency: There are evaluators who limit their evaluation to the utterance in line 91, "..." 4. Voice expression: Major tendency: The evaluation is mainly based on the statement "..." on line 201. 5. Language: Major tendency: The evaluation is mainly based on the statement "..." on line 123.

[0046] The generation AI model L can determine the degree of bias in the reasons for the first evaluations of the subject's voice in multiple pieces of first evaluation information D1. For example, after the subject uses incorrect language in only one part of the voice information, the evaluator may make a judgment based on that one part in the first evaluations of all of the first evaluation information D1. In this case, the quality of the language used is reflected in, for example, the entire voice information, so it can be determined that the bias in the reasons for the first evaluations of the multiple pieces of first evaluation information D1 is large. Furthermore, it can be determined that the evaluator makes a judgment based on various parts of the voice information in the first evaluations of all of the first evaluation information D1. In this case, it can be determined that the bias in the reasons for the first evaluations of the multiple pieces of first evaluation information D1 is small.

[0047] On the other hand, there may be cases where the subject of evaluation confirms the purpose of the audio information only at the beginning, and then all the evaluators make their judgments based on that one part. In this case, the quality of the purpose of the audio information confirmation is expressed locally in the audio information, for example, so it can be determined that there is little bias in the reasons for the first evaluations of the multiple pieces of first evaluation information D1.

[0048] Next, the processing unit 30 generates second evaluation information D5 based on the first evaluation information D1 and the trend information D2 (see FIGS. 1(c) and 1(d)). Specifically, the processing unit 30 executes processing on the voice information, the first evaluation information D1, and the trend information D2 to obtain the second evaluation information D5 using the generation AI model L so that the trend of the reasons for the second evaluation in the second evaluation information D5 includes the trend indicated by the trend information D2 and other trends different from the trend.

[0049] More specifically, the processing unit 30 generates third evaluation information D4 based on the trend information D2 (see FIG. 1C). The processing unit 30 processes the voice information and the trend information D2 to obtain the third evaluation information D4 using the generation AI model L so that the reason for the third evaluation in the third evaluation information D4 has a different tendency from the tendency indicated by the trend information D2. The third evaluation information D4 is information indicating the third evaluation of the voice of the person being evaluated and the reason for the third evaluation (details will be described later).

[0050] For example, the processing unit 30 generates input information based on the tendency information D2 and text information and includes the generated input information in a prompt (hereinafter referred to as a second general-purpose prompt) stored in the speech evaluation device 10, thereby generating a prompt D3. FIG. 6 is a diagram showing an example of the prompt D3. In the example shown in FIG. 6, the second general-purpose prompt is an instruction including instructions D31, D32, and D33. The processing unit 30 inputs input information including the tendency information D2 and text information into the input fields of the second general-purpose prompt, and generates instructions D34 and D35, thereby generating the prompt D3. In the second general-purpose prompt, character strings indicating that the tendency information D2 and text information are to be input are provided after the instructions D32 and D33, respectively.

[0051] The instruction D31 of the prompt D3 includes information to generate third evaluation information D4 based on the text information and trend information D2. In the example shown in FIG. 6, the information is as follows: "#Instruction #The interaction history is a conversation between a staff member and a customer. Each line includes the line number, speaker ID, text of the speech recognition result, emotion recognition result, and speaking rate information. #The evaluation trend shows the tendency of the evaluator's evaluation of the staff member's interaction. #Referring to the interaction history and #the evaluation trend, please use a three-point scale to determine whether the staff member has achieved the #check items while interacting with customers, and please also tell us the reason. When doing so, please try to evaluate from a perspective that differs as much as possible from the #evaluation trend."

[0052] The above section includes information explaining the text information. The information is the following section: "# Response history is a conversation between a staff member and a customer. Each line includes a line number, a speaker ID, a text of the speech recognition result, an emotion recognition result, and speaking rate information." The line number, speaker ID, text of the speech recognition result, an emotion recognition result, and speaking rate information are generated from the speech information using known methods.

[0053] The above section includes content explaining the trend information D2. This information is the section "#Evaluation trends show the trends in the evaluations made by evaluators regarding the staff's responses."

[0054] The above section includes information to the effect that third evaluation information D4 will be generated based on the text information and trend information D2, and content to the effect that third evaluation information D4 will be generated by evaluating the text information in a manner different from the tendency of the evaluation reasons (perspectives) in the trend information D2. The information in question is the section that reads, "Referring to the # service history and # evaluation trends, please judge on a three-point scale whether the staff member has achieved the # check items while serving customers, and please also tell us the reason. When doing so, please try to evaluate from a perspective different from the # evaluation trends as much as possible."

[0055] The instruction sentence D32 of the prompt D3 includes content that explains each of the multiple items contained in the third evaluation information D4. In the example shown in Fig. 6, the information is the following: "#Check items 1. Confirmation of purpose: The person repeats the customer's purpose. If it is not clear, the person asks questions to clarify. 2. Ability to explain: The person explains concisely in words that the customer can understand. 3. Pause and speed of speech: The person speaks at a pace and speed that suits the customer. 4. Voice expression: The person speaks with a voice expression appropriate to the situation. 5. Language: The person speaks using business terms and honorific language."

[0056] The instruction sentence D33 of the prompt D3 includes information indicating that the third evaluation information D4 includes, for each item, an evaluation of the speech of the person being evaluated and information indicating the reason for the evaluation. In the example shown in Fig. 6, the information is "#Output format 1. Confirmation of purpose: [A / B / C] Reason: [Enter reason] 2. Ability to explain: [A / B / C] Reason: [Enter reason] 3. Pause and speed of speech: [A / B / C] Reason: [Enter reason] 4. Facial expression of voice: [A / B / C] Reason: [Enter reason] 5. Use of language: [A / B / C] Reason: [Enter reason]".

[0057] The instruction D34 of the prompt D3 includes content indicating the tendency information D2. The instruction D34 includes content indicating a plurality of items, the degree of bias in the reasons for evaluation for each item, and the tendency of the reasons for evaluation corresponding to each item. In the example shown in FIG. 6 , the information is the following: "1. Confirmation of the purpose: minor tendency: evaluation is conducted based on the repetition of the requirements in "..." on line 5. 2. Explanation ability: medium tendency: evaluation is conducted based on the length of the explanation and evaluation based on rephrasing. 3. Speech timing and speed: medium tendency: some evaluators limit their evaluation to the statement in "..." on line 91. 4. Voice expression: major tendency: evaluation is conducted mainly based on the statement in "..." on line 201. 5. Language: major tendency: evaluation is conducted mainly based on the statement in "..." on line 123."

[0058] The instruction D65 of the tendency evaluation prompt D6 includes the content of the text information, which is the "# response history..." portion.

[0059] The processing unit 30 inputs the generated prompt D3 into the generation AI model L to obtain third evaluation information D4 generated by the generation AI model L. The third evaluation information D4 is information indicating a third evaluation of the voice of the person being evaluated and the reason for the third evaluation for each item. The multiple items are set in advance by a user or the like depending on the type of voice of the person being evaluated. For example, the multiple items are items related to the satisfaction level of customers who called the call center. The multiple items related to the third evaluation are typically the same as the multiple items related to the first evaluation. The third evaluation is a third evaluation result indicating an evaluation of the voice of the person being evaluated using one of multiple stages. The stage related to the third evaluation may be the same as the stage related to the first evaluation. The reason for the third evaluation is, for example, a sentence explaining the reason. The third evaluation information D4 may have the same format as the first evaluation information D1.

[0060] FIG. 7 is a diagram showing an example of the third evaluation information D4. In the example shown in FIG. 7, the third evaluation information D4 indicates a third evaluation based on the tendency information D2 and text information and the reason for the third evaluation for each item. The information is the following part: "1. Confirmation of the purpose: A Reason: The purpose was repeated in "..." on line 5. 2. Explanation ability: A Reason: The explanation was not tailored to the other person's reaction, such as by rephrasing the term in "..." on line 37. 3. Pause and speed of speech: B Reason: The speaker spoke at a faster pace than the other person on average. 4. Voice expression: A Reason: The speaker apologized in a tone of voice that sounded sad in "..." on line 137. 5. Language: A Reason: The only inappropriate language was in "..." on line 123."

[0061] The processing unit 30 generates second evaluation information D5 based on the voice information, the first evaluation information D1, and the third evaluation information D4 (see FIG. 1(d)). Specifically, the processing unit 30 processes the voice information, the first evaluation information D1, and the third evaluation information D4 to obtain the second evaluation information D5 using the generation AI model L so that the tendency of the reasons for the second evaluation in the second evaluation information D5 includes the tendency indicated by the tendency information D2 and another tendency different from the tendency. The processing unit 30 generates input information based on the multiple pieces of first evaluation information D1 and third evaluation information D4, and includes the input information in a prompt (hereinafter referred to as a third general-purpose prompt) stored in the voice evaluation device 10, thereby generating a prompt D7.

[0062] 8 is a diagram showing an example of the prompt D7. In the example shown in FIG. 8, the third general-purpose prompt is a command statement including command statements D71, D72, and D73. The processing unit 30 inputs input information including a plurality of pieces of first evaluation information D1 and a plurality of pieces of third evaluation information D4 into the input fields of the third general-purpose prompt, and generates command statement D74, thereby generating the prompt D7. In the third general-purpose prompt, character strings indicating that the plurality of pieces of first evaluation information D1 and the third evaluation information D4 are to be input are provided after command statement D72, respectively.

[0063] The instruction D71 of the prompt D7 includes information to generate second evaluation information D5 based on the plurality of first evaluation information D1 and third evaluation information D4, and information to generate second evaluation information D5 that satisfies the conditions described in the instruction D72. The information is the part "#Instruction #Evaluation result by evaluator is the evaluation result by the evaluator of the staff's response. Please refer to the evaluation results by the plurality of #evaluators and create three patterns of composite evaluation results that satisfy the #composite result creation conditions."

[0064] The instruction statement D72 of the prompt D7 includes the content indicated by the conditions for generating the second evaluation information D5. The instruction statement D72 includes three conditions. In the example shown in FIG. 8 , the information is the following: "#Conditions for creating composite results <Pattern 1> - For the judgment result, the content of the majority should be adopted. - For the reason, the content of the evaluation result with the two most different perspectives should be reflected. <Pattern 2> - For the judgment result, the content of the majority should be adopted. - For the reason, the perspectives expressed in all evaluation results should be covered. <Pattern 3> - For the judgment result, the one with a clear reason for reaching the judgment should be adopted..."

[0065] The above section includes adopting the majority evaluation from the multiple first evaluation results and the third evaluation results as the second evaluation of the speech of the subject. For example, the most common level among the multiple levels of "A," "B," and "C" in the first evaluation of the multiple first evaluation results and the third evaluation of the third evaluation results becomes the second evaluation. Furthermore, the above section includes extracting two evaluation results from the multiple first evaluation results and the multiple third evaluation results that have different reasons, and adopting the respective reasons for the two evaluation results as the reason for the second evaluation. This information is the section that reads, "<Pattern 1> - For the judgment result, the content of the majority is adopted - For the reason, the content of the evaluation results with the two most different perspectives is reflected."

[0066] The second pattern of instruction sentence D72 includes adopting the majority evaluation among the multiple first evaluation results and the third evaluation results as the second evaluation of the speech of the subject, and adopting all of the reasons for the multiple first evaluation results and the reasons for the third evaluation results as the reason for the second evaluation. This information is the part that reads "<Pattern 2> - The majority content should be adopted as the judgment result - The reasons should cover the perspectives that appear in all evaluation results."

[0067] As shown in the first and second patterns of instruction statement D72, for example, the processing unit 30 performs processing on the audio information, the first evaluation information D1, and the third evaluation information D4 to obtain second evaluation information D5 using the generation AI model L, so as to determine the second evaluation result based on the number of times each stage is applied in the first evaluation result and the third evaluation result.

[0068] The third pattern of instruction sentence D72 includes adopting the first or third rating as the second rating when the reason for the first or third rating of the voice of the subject is clear in the multiple first and third evaluation results. The reason for the first or third rating being clear means, for example, that the reason for the first or third rating is explained in a way that allows a person to clearly understand the reason. This information is the part that reads, "<Pattern 3> - Regarding the judgment result, adopt the one with a clear reason for reaching the judgment..."

[0069] The instruction sentence D73 of the prompt D7 includes information indicating that the second evaluation information D5 includes, for each item, a second evaluation of the speech of the person being evaluated and information indicating the reason for the second evaluation. In the example shown in Fig. 8, the information is "#Output format 1. Confirmation of purpose: [A / B / C] Reason: [Enter reason] 2. Ability to explain: [A / B / C] Reason: [Enter reason] 3. Pause and speed of speech: [A / B / C] Reason: [Enter reason] 4. Facial expression of voice: [A / B / C] Reason: [Enter reason] 5. Use of language: [A / B / C] Reason: [Enter reason]".

[0070] The instruction sentence D74 of the prompt D7 includes information indicating a plurality of pieces of first evaluation information D1 and second evaluation information D5. In the example shown in FIG. 8 , the information is the following: " #Evaluation results by evaluators <Evaluation result of first evaluator> 1. Confirmation of the subject matter: A Comment: The subject matter was repeated in "..." on line 5 and "..." on line 38. The subject matter was not repeated in "..." on line 59, but the subject matter was rephrased and confirmed in "..." on line 66. 2. Explanation ability: B Comment: While various paraphrases were used to communicate in words that were easy for customers to understand, there was a tendency for sentences to be lengthy throughout. 3. Speech timing and speed: A Comment: No particular problems were found. 4. Voice expression: A Comment: The speaker spoke with an appropriate voice and expression. 5. Language: B Comment: The speaker used the expression "Was it okay?" <Evaluation result of second evaluator> ... <Evaluation result of third evaluator> ... "

[0071] The processing unit 30 inputs the generated prompt D7 into the generation AI model L, thereby obtaining second evaluation information D5 generated by the generation AI model L. The processing unit 30 presents the second evaluation information D5 to the person being evaluated by outputting the second evaluation information D5 to a display or the like. The second evaluation information D5 includes, for each item, a second evaluation of the person being evaluated's voice based on the plurality of pieces of first evaluation information D1 and third evaluation information D4, and information indicating the reason for the second evaluation. FIG. 9 is a diagram showing an example of the second evaluation information D5. In the example shown in FIG. 9, the second evaluation information D5 indicates, for each item, the second evaluation based on the plurality of pieces of first evaluation information D1 and third evaluation information D4, and the reason for the second evaluation, in accordance with the conditions of each of the first to third patterns. The information is as follows: <Pattern 1> 1. Confirmation of the purpose: A Reason: The purpose was repeated in the fifth line "...". 2. Explanation ability: B Reason: The sentence was long and redundant. 3. Pause and speed of speech: A Reason: There was no big change from beginning to end. 4. Voice expression: A Reason: There was no tone of voice that did not fit the situation. 5. Diction: B Reason: The expression "..." on line 123 was inappropriate. <Pattern 2> This is the "..." part.

[0072] Next, a voice evaluation method, which is a process executed by the voice evaluation device 10 according to this embodiment, will be described with reference to the flowchart shown in FIG.

[0073] First, the voice evaluation device 10 acquires voice information and first evaluation information D1 (step S1: acquisition step) (see FIG. 1A). Then, the voice evaluation device 10 generates text information based on the voice information (step S2).

[0074] The voice evaluation device 10 generates tendency information D2 based on the text information and the first evaluation information D1 (step S3) (see FIG. 1(b)). Specifically, the voice evaluation device 10 executes processing on the text information and the first evaluation information D1 to obtain tendency information D2 based on the text information and the first evaluation information D1 using the generation AI model L.

[0075] Finally, as shown in steps S4 and S5, the voice evaluation device 10 performs processing on the voice information, the first evaluation information D1, and the trend information D2 to obtain the second evaluation information D5 using the generation AI model L so that the trend of the reasons for the second evaluation in the second evaluation information D5 includes the trend indicated by the trend information D2 and other trends different from that trend (see Figures 1(c) and 1(d)).

[0076] Specifically, the voice evaluation device 10 obtains third evaluation information D4 based on the text information and tendency information D2 (step S4) (see FIG. 1(c)). More specifically, the voice evaluation device 10 processes the voice information and tendency information D2 to obtain third evaluation information D4 using a generative AI model so that the reason for the third evaluation has another tendency.

[0077] The audio evaluation device 10 obtains second evaluation information based on the first evaluation information D1 and the third evaluation information D4 (step S5) (see FIG. 1(d)). More specifically, the audio evaluation device 10 processes the audio information, the first evaluation information D1, and the third evaluation information D4 using the generation AI model L to obtain the second evaluation information D5 so that the tendency of the reasons for the second evaluation in the second evaluation information D5 includes the tendency indicated by the tendency information D2 and the other tendency.

[0078] In this way, the audio evaluation device 10 executes processing on the audio information and the first evaluation information D1 to obtain the second evaluation information D5 based on the audio information and the first evaluation information D1 using the generation AI model L (steps S2 to S5: processing steps).

[0079] Next, the effects of the voice evaluation device 10 and voice evaluation method according to this embodiment will be described. According to this embodiment, processing is performed on the voice information and the first evaluation information D1 using a generation AI model L generated by machine learning to obtain second evaluation information D5 indicating a second evaluation of the voice based on the voice information and first evaluation information D1 indicating the first evaluation of the voice of the person being evaluated and the reason for the first evaluation. With this configuration, the second evaluation information D5 is generated based at least on the reason for the first evaluation in the first evaluation information D1. This makes it possible to suppress bias in the reason for the second evaluation in the second evaluation information D5 compared to bias in the reason for the first evaluation (e.g., viewpoint) in the first evaluation information D1, thereby ensuring the quality of the second evaluation information D5. As a result, the quality of the evaluation results of the voice of the person being evaluated can be guaranteed. Furthermore, bias in the reason for the second evaluation in the second evaluation information D5 can be suppressed without increasing the number of pieces of first evaluation information D1. This makes it possible to ensure the quality of the second evaluation information D5 while reducing the personnel costs required for increasing the number of pieces of first evaluation information D1. As a result, it is possible to reduce personnel costs while ensuring the quality of the evaluation results of the speech of the person being evaluated.

[0080] As in the above-described embodiment, the processing unit 30 may generate a prompt D7 to be input to the generating AI model L based on the voice information and the first evaluation information D1. This configuration enables the generating AI model L to generate the second evaluation information D5. Furthermore, by adjusting the description of the prompt D7, the desired second evaluation information D5 can be reliably obtained. However, the processing unit 30 may also obtain the second evaluation information D5 by the generating AI model L by a method other than generating the prompt D7.

[0081] As in the above-described embodiment, the processing unit 30 may input input information based on the voice information and the first evaluation information D1 to the generation AI model L to obtain the second evaluation information D5. With this configuration, the generation AI model L can generate the second evaluation information D5.

[0082] As in the above-described embodiment, the second evaluation information D5 may be information indicating the reason for the second evaluation of the voice of the subject of evaluation. In this case, it is possible to obtain second evaluation information D5 that covers all potential reasons for the second evaluation, thereby obtaining second evaluation information D5 in which the degree of bias in the reasons for the second evaluation is suppressed. This makes it possible to improve, for example, the quality of feedback to the subject of evaluation (operator) based on the second evaluation information D5. However, the second evaluation information D5 may be information indicating only the second evaluation.

[0083] As in the above-described embodiment, the processing unit 30 may process the audio information and the first evaluation information D1 to obtain, using the generation AI model L, trend information D2 indicating the tendency of the reasons for the first evaluation of the first evaluation information D1 based on the audio information and the first evaluation information D1. The processing unit 30 may also process the audio information, the first evaluation information D1, and the trend information D2 to obtain, using the generation AI model L, the second evaluation information D5 so that the tendencies of the reasons for the second evaluation of the second evaluation information D5 include the tendency indicated by the trend information D2 and other trends different from the tendency. According to this configuration, the second evaluation information D5 can be generated based on the audio information, the first evaluation information D1, and the trend information D2 so that the reasons for the second evaluation include the tendency indicated by the trend information D2 and other trends different from the tendency. This allows the generation of second evaluation information D5 that more reliably covers potential reasons for the second evaluation indicated by the second evaluation information D5. However, the audio evaluation device 10 may generate the second evaluation information D5 based on multiple pieces of first evaluation information D1 and audio information without generating the trend information D2.

[0084] As in the above-described embodiment, the processing unit 30 may process the voice information and the trend information D2 to obtain, using the generation AI model L, third evaluation information D4 indicating a third evaluation of the voice of the subject and the reason for the third evaluation, such that the reason for the third evaluation has other tendencies. Alternatively, the processing unit 30 may process the voice information, the first evaluation information D1, and the third evaluation information D4 to obtain, using the generation AI model L, second evaluation information D5 such that the tendencies of the reasons for the second evaluation in the second evaluation information D5 include the tendencies indicated by the trend information D2 and other tendencies. This configuration makes it possible to generate second evaluation information D5 that more reliably covers potential reasons for the second evaluation. However, the voice evaluation device 10 may generate second evaluation information D5 based on multiple pieces of first evaluation information D1 and voice information without generating third evaluation information D4.

[0085] As in the above-described embodiment, the first evaluation may be a first evaluation result indicating an evaluation of the subject's voice using one of a plurality of stages indicating the quality of the evaluation of the subject's voice, the second evaluation may be a second evaluation result indicating an evaluation of the subject's voice using one of a plurality of stages, and the third evaluation may be a third evaluation result indicating an evaluation of the subject's voice using one of a plurality of stages. The processing unit 30 may execute processing on the voice information, the first evaluation information D1, and the third evaluation information D4 using the generation AI model L to obtain second evaluation information D5 so as to determine the second evaluation result based on the number of times each stage is applied in at least one first evaluation result and at least one third evaluation result. This configuration reliably ensures the quality of the evaluation result of the subject's voice.

[0086] The configuration of the sound evaluation device 10 described in the above embodiment is an example, and various modifications can be made. Representative modifications will be described below.

[0087] In the above embodiment, the processing unit 30 generates the tendency information D2 based on the audio information and the first evaluation information D1, generates the third evaluation information D4 based on the audio information, the first evaluation information D1, and the tendency information D2, and calculates the second evaluation information D5 based on the audio information, the first evaluation information D1, and the third evaluation information D4, but is not limited to this. For example, the processing unit 30 may generate the second evaluation information D5 based on the audio information and the first evaluation information D1 without generating the tendency information D2 and the third evaluation information D4.

[0088] 11 and 12 are diagrams showing prompts D103 and D203, respectively, for generating second evaluation information D105 in a modified example. Fig. 13 is a diagram showing an example of second evaluation information D105 in a modified example. In the example shown in Fig. 11, the second general prompt is a command statement including command statements D131, D32, and D33.

[0089] Unlike the above embodiment, the instruction sentence D31 of the prompt D103 includes information to evaluate the text information based on tendencies that include tendencies different from the tendencies of the reasons for the first evaluations in the plurality of pieces of first evaluation information D1, and tendencies different from the tendencies of the reasons for the first evaluations. In the example shown in Figure 14, the information is the section that reads, "At this time, please quote as much as possible comments from the response history that are not mentioned in the evaluation results by the #evaluator, and output a comprehensive evaluation result that also takes into account the evaluation results by the #evaluator."

[0090] As in the above embodiment, the instruction statement D32 of the prompt D103 includes information explaining each of the multiple items contained in the second evaluation information D5. As in the above embodiment, the instruction statement D33 of the prompt D103 includes content to the effect that the second evaluation information D105 includes, for each item, an evaluation of the voice of the person being evaluated and information indicating the reason for the evaluation. The instruction statement D134 of the prompt D103 includes content indicating multiple pieces of first evaluation information D1. In the example shown in Figure 14, an instruction statement similar to the instruction statement D64 shown in Figure 4 is written as the instruction statement D134. As in the above embodiment, the instruction statement D35 of the prompt D103 includes the content of text information. The information in question is the "# response history..." portion.

[0091] In the example shown in FIG. 12 , the second general-purpose prompt is an instruction including instructions D231, D32, and D33. Unlike instruction D131 of prompt D103, instruction D231 of prompt D203 includes information indicating that the text information will be evaluated based on trends in the reasons for the first evaluations in the plurality of pieces of first evaluation information D1, as well as trends different from the trends in the reasons for the first evaluations. For example, instruction D231 includes content indicating that the voice of the person being evaluated will be evaluated based on parts of the voice of the person being evaluated that are not included as reasons for the first evaluations in the plurality of pieces of first evaluation information D1. In the example shown in FIG. 14 , the information in question is the section that reads, "At this time, please quote, as much as possible, statements from the response history that are not mentioned in the evaluation results by the evaluator."

[0092] In the above embodiment and modified example, the processing unit 30 inputs a prompt D7 to be input to the generated AI model L based on the voice information and the first evaluation information D1 to the generated AI model L and obtains the second evaluation information D5. However, it is not necessary to obtain the second evaluation information D5. In this case, the processing unit 30 generates a prompt D7 to be input to the generated AI model L based on the voice information and the first evaluation information D1, and outputs the generated prompt D7 to another device (e.g., terminal A). The prompt D7 is transmitted to the server device 100 by the other device and input to the generated AI model L. The server device 100 transmits the second evaluation information D5 output from the generated AI model L to terminal A. The process of causing the generated AI model to generate the second evaluation information D5 using the prompt D7 in this manner may be performed by the other device.

[0093] In the above embodiment, the voice evaluation device 10 is a server device separate from terminal A, but this is not limited to this. For example, the voice evaluation device 10 may be part of terminal A. As an example, the RAG system configured by the acquisition unit 20 and the processing unit 30 may be implemented in terminal A. In this case, the knowledge database searched by the RAG system may be implemented inside terminal A, may be implemented in the voice evaluation device 10, or may be implemented in another device on a network (e.g., the cloud).

[0094] In the above embodiment, the information processing system 1 includes a terminal A, a voice evaluation device 10, and a server device 100, but is not limited thereto. For example, the information processing system 1 of a modified example may include only a voice evaluation device 10 according to the modified example. In this case, the voice evaluation device 10 is a terminal held by a user, such as a device used by a user who wishes to obtain various types of content based on input information using interactive AI. The voice evaluation device 10 includes a RAG system consisting of an acquisition unit 20 and a processing unit 30, and a generation AI model L. The configuration of the above modified example can be realized by installing an application that executes the functions of the generation AI model L in the voice evaluation device 10. In this way, the generation AI model L may be implemented on a network (e.g., cloud) other than the server device 100. Note that in the voice evaluation device 10 of the modified example, the knowledge database searched by the RAG system may be implemented within the voice evaluation device 10 or on another device on the network (e.g., cloud).

[0095] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are directly or indirectly connected (e.g., wired, wireless, etc.) and these multiple devices. The functional block may also be realized by combining software with the single device or multiple devices.

[0096] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0097] For example, the audio evaluation device 10 according to an embodiment of the present disclosure may function as a computer that performs information processing according to the present disclosure. Fig. 14 is a diagram showing an example of the hardware configuration of the audio evaluation device 10 according to an embodiment of the present disclosure. The audio evaluation device 10 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, etc.

[0098] In the following description, the term "apparatus" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the sound evaluation apparatus 10 may be configured to include one or more of the apparatuses shown in the drawings, or may be configured to exclude some of the apparatuses.

[0099] Each function of the audio evaluation device 10 is realized by loading predetermined software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.

[0100] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control device, an arithmetic unit, a register, etc. For example, the processing unit 30 of the sound evaluation device 10 may be realized by the processor 1001.

[0101] The processor 1001 also reads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the processing unit 30 of the audio evaluation device 10 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be made for other functional blocks. While the above-described various processes have been described as being executed by one processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0102] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for performing information processing according to an embodiment of the present disclosure.

[0103] The storage 1003 is a computer-readable recording medium, and may be composed of at least one of an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. The storage 1003 may also be called an auxiliary storage device. The storage medium provided in the voice evaluation device 10 may be, for example, a database, a server, or other appropriate medium including at least one of the memory 1002 and the storage 1003.

[0104] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also called, for example, a network device, a network controller, a network card, or a communication module.

[0105] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. The input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).

[0106] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0107] The sound evaluation device 10 may also be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0108] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0109] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0110] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0111] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0112] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0113] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0114] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0115] As used in this disclosure, the terms "system" and "network" are used interchangeably.

[0116] Furthermore, the information, parameters, etc. described in this disclosure may be expressed using absolute values, may be expressed using relative values ​​from a predetermined value, or may be expressed using other corresponding information.

[0117] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0118] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0119] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0120] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0121] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.

[0122] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0123] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0124] The speech evaluation device and speech evaluation method disclosed herein have the following configuration. [1] A speech evaluation device comprising: an acquisition unit that acquires speech information related to the speech of a subject person, as well as first evaluation information indicating a first evaluation of the speech and a reason for the first evaluation; and a processing unit that executes processing on the speech information and the first evaluation information to obtain, using a generative AI model generated by machine learning, second evaluation information indicating a second evaluation of the speech based on the speech information and the first evaluation information. [2] The speech evaluation device described in [1], wherein the processing unit generates a prompt to be input to the generative AI model based on the speech information and the first evaluation information. [3] The speech evaluation device described in [1], wherein the processing unit inputs input information based on the speech information and the first evaluation information into the generative AI model to obtain the second evaluation information. [4] The speech evaluation device described in any of [1] to [3], wherein the second evaluation information is information further indicating a reason for the second evaluation of the speech. [5] The voice evaluation device described in any of [1] to [4], wherein the processing unit processes the voice information and the first evaluation information to obtain, by the generative AI model, trend information indicating a trend of the reason for the first evaluation of the first evaluation information based on the voice information and the first evaluation information, and processes the voice information, the first evaluation information, and the trend information to obtain, by the generative AI model, the second evaluation information so that the trend of the reason for the second evaluation of the second evaluation information includes the trend indicated by the trend information and another trend different from the trend. [6] The processing unit processes the voice information and the tendency information to obtain, by the generative AI model, a third evaluation of the voice and third evaluation information indicating a reason for the third evaluation, such that the reason for the third evaluation has the other tendency; and processes the voice information, the first evaluation information, and the third evaluation information to obtain, by the generative AI model, the second evaluation information, such that the tendency of the reason for the second evaluation in the second evaluation information includes the tendency and the other tendency indicated by the tendency information. The voice evaluation device described in [5].[7] The voice evaluation device according to [6], wherein the first evaluation is a first evaluation result indicating an evaluation of the voice using one of a plurality of stages indicating a good evaluation of the voice, the second evaluation is a second evaluation result indicating an evaluation of the voice using one of the plurality of stages, and the third evaluation is a third evaluation result indicating an evaluation of the voice using one of the plurality of stages, and the processing unit processes the voice information, the first evaluation information, and the third evaluation information to obtain the second evaluation information using the generative AI model so as to determine the second evaluation result based on the number of times each of the stages has been applied in at least one or more of the first evaluation results and at least one or more of the third evaluation results. [8] A voice evaluation method comprising: an acquisition step of acquiring voice information related to the voice of a subject to evaluation, and first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation; and a processing step of processing the voice information and the first evaluation information to obtain second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information using a generative AI model generated by machine learning.

[0125] 10...voice evaluation device, 20...acquisition unit, 30...processing unit, L...generative AI model, D1...first evaluation information, D2...trend information, D4...third evaluation information, D5...second evaluation information, D7...prompt, 1001...processor, 1002...memory, 1003...storage, 1004...communication device, 1005...input device, 1006...output device.

Claims

1. A voice evaluation device comprising: an acquisition unit that acquires voice information regarding the voice of a person to be evaluated, as well as first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation; and a processing unit that executes processing on the voice information and the first evaluation information to obtain second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information using a generative AI model generated by machine learning.

2. The voice evaluation device according to claim 1, wherein the processing unit generates a prompt to be input to the generative AI model based on the voice information and the first evaluation information.

3. The voice evaluation device according to claim 1, wherein the processing unit inputs input information based on the voice information and the first evaluation information into the generative AI model to obtain the second evaluation information.

4. The audio evaluation device according to claim 1, wherein the second evaluation information is information that further indicates the reason for the second evaluation of the audio.

5. The voice evaluation device described in claim 4, wherein the processing unit processes the voice information and the first evaluation information to obtain, by the generative AI model, trend information indicating a trend of the reason for the first evaluation of the first evaluation information based on the voice information and the first evaluation information, and processes the voice information, the first evaluation information, and the trend information to obtain, by the generative AI model, the second evaluation information so that the trend of the reason for the second evaluation of the second evaluation information includes the trend indicated by the trend information and another trend different from the trend.

6. The voice evaluation device described in claim 5, wherein the processing unit processes the voice information and the tendency information to obtain, by the generative AI model, third evaluation information indicating a third evaluation of the voice and a reason for the third evaluation, such that the reason for the third evaluation has the other tendency; and processes the voice information, the first evaluation information, and the third evaluation information to obtain, by the generative AI model, the second evaluation information, such that the tendency of the reason for the second evaluation in the second evaluation information includes the tendency and the other tendency indicated by the tendency information.

7. The voice evaluation device of claim 6, wherein the first evaluation is a first evaluation result indicating an evaluation of the voice using one of a plurality of stages indicating a good evaluation of the voice, the second evaluation is a second evaluation result indicating an evaluation of the voice using one of the plurality of stages, and the third evaluation is a third evaluation result indicating an evaluation of the voice using one of the plurality of stages, and the processing unit performs processing on the voice information, the first evaluation information, and the third evaluation information to obtain the second evaluation information using the generative AI model, so as to determine the second evaluation result based on the number of times each of the stages is applied in the first evaluation result and the third evaluation result.

8. A voice evaluation method comprising: an acquisition step of acquiring voice information regarding the voice of a person to be evaluated, as well as first evaluation information indicating a first evaluation of the voice and a reason for the first evaluation; and a processing step of executing processing on the voice information and the first evaluation information to obtain second evaluation information indicating a second evaluation of the voice based on the voice information and the first evaluation information using a generative AI model generated by machine learning.

Citation Information

Patent Citations

  • Automatic scoring device for dialog between operator and customer, and operation method for the same

    JP2014123813A

  • Language production ability evaluation system, language production ability evaluation program, and language production ability evaluation method

    JP7521860B1