Conversation evaluation program, device and method
The conversation evaluation program objectively assesses communication states by estimating speaker utterances and changes, enhancing the accuracy of conversation evaluations.
Patent Information
- Application Number
- JP2023091036
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-06-01
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2043-06-01
AI Technical Summary
Existing technologies fail to provide an objective evaluation of the state of communication between multiple speakers, as they rely solely on the perspective of one speaker, leading to potential unfair assessments.
A conversation evaluation program that estimates the start and end times of utterances by dominant speakers, identifies speaker changes, and evaluates the conversation state based on dialogue information including overlapping and silent sections, using a trained detection model to generate an objective evaluation.
Enables an objective evaluation of communication states between multiple speakers, identifying intimidating speaker changes and activity levels, improving the accuracy of conversation assessments.
Smart Images

Figure 00000013_0000 
Figure 00000013_0001 
Figure 00000014_0000
Abstract
Description
[Technical Field]
[0001] SUMMARY OF THE INVENTION Embodiments of the present invention relate to speech assessment programs, devices and methods. [Background technology]
[0002] Generally, to improve labor productivity, companies need to improve the engagement of each employee with the company. For this reason, companies need to evaluate the state (especially the health) of communication between employees.
[0003] For example, there is a technology that evaluates the state of communication between two speakers in a conversation between a first speaker and a second speaker, using the length of the overlapping section where the two speakers' speech sections overlap. This technology considers the overlapping section to be the time from when the second speaker starts speaking while the first speaker is speaking to when the first speaker finishes speaking, and evaluates the second speaker's impression of the first speaker as "average" or "bad" depending on the length of this overlapping section.
[0004] However, because the above technology evaluates the impression of the first speaker from the perspective of the second speaker, it is not possible to objectively evaluate the impression of the first speaker independently of the second speaker's perspective. For example, the impression of the first speaker may be unfairly evaluated as "bad" even though the second speaker who interrupted the first speaker's speech was actually bad. Therefore, there is a demand for an appropriate evaluation of the state of communication between multiple speakers. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent No. 6524674 [Non-patent literature]
[0006] [Non-Patent Document 1] Nobutaka Ito, Christopher Schymura, Shoko Araki and Tomohiro Nakatani, "Noisy cGMM: Complex Gaussian Mixture Model with Non-Sparse Noise Model for Joint Source Separation and Denoising", 2018 26th European Signal Processing Conference (EUSIPCO), Rome, Italy, 2018, pp. 1662-1666 [Non-patent document 2] Jongseo Sohn, Nam Soo Kim and Wonyong Sung, "A statistical model-based voice activity detection", IEEE Signal Processing Letters, 1999, Vol. 6, No. 1, pp. 1-3 Summary of the Invention [Problem to be solved by the invention]
[0007] The problem to be solved by the present invention is to appropriately evaluate the state of communication between multiple speakers. [Means for solving the problem]
[0008] A conversation evaluation program according to an embodiment causes a computer to implement an estimation function, an identification function, and an evaluation function. The estimation function estimates the start and end times of utterances by each dominant speaker based on audio data related to a conversation including utterances by multiple speakers. The identification function identifies a timing for each dominant speaker to take a turn based on the estimated start and end times. The evaluation function evaluates the state of the conversation based on dialogue information before and after the identified timing for taking a turn. The dialogue information includes at least one of the lengths of overlapping sections where the speech sections of the dominant speakers overlap and the lengths of silent sections where the speech sections of the dominant speakers do not overlap, as well as the lengths of the speech sections of the dominant speakers. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing an example of the functional configuration of a conversation evaluation device according to an embodiment of the present invention. [Figure 2] FIG. 1 is a block diagram showing an example of the hardware configuration of a conversation evaluation device according to an embodiment of the present invention. [Figure 3] 1 is a flowchart showing an example of the overall operation of the conversation evaluation device according to the present embodiment. [Figure 4] 10 is a flowchart showing a method for estimating the speech of a main speaker according to the present embodiment. [Figure 5] FIG. 1 is a diagram showing an example of conversation analysis according to the present embodiment. [Figure 6] 1 is a flowchart showing a method for evaluating a conversation state according to the present embodiment. [Figure 7] FIG. 4 is a diagram showing an example of dialogue information according to the embodiment. [Figure 8] FIG. 10 is a diagram showing an example of displaying an evaluation result according to the embodiment. [Figure 9] FIG. 10 is a diagram showing an example of the estimation accuracy of conversation analysis according to the conventional method and the proposed method. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, a conversation evaluation program, an apparatus, and a method according to the embodiments will be described with reference to the drawings. In the following embodiments, parts with the same reference numerals perform similar operations, and redundant explanations will be omitted as appropriate.
[0011] The following are definitions of each term used in this embodiment. (1) "Utterance" refers to a speech interval (voice interval) sandwiched between two silent intervals of 0.2 seconds or more. (2) "Dominant speaker" refers to a speaker in a conversation. (3) "Minor speaker" refers to a listener to the dominant speaker. The speech content of a minor speaker often includes backchannel responses or fragmentary repetitions of the dominant speaker's utterance. (4) "Substitution of dominant speaker" refers to a substitution between two dominant speakers. (5) "Dominant speaker's turn" refers to the set of utterances made by the dominant speaker from the time the dominant speaker takes over until the time the dominant speaker takes over. (6) "Length of dominant speaker's speech interval" refers to the total length of all speech intervals in one dominant speaker's turn.
[0012] 1 is a block diagram showing an example of the functional configuration of a conversation evaluation device 1 according to this embodiment. The conversation evaluation device 1 is a device that evaluates the state of a conversation between multiple speakers. The conversation evaluation device 1 includes an acquisition unit 111, an estimation unit 112, an identification unit 113, and an evaluation unit 114.
[0013] The acquisition unit 111 acquires various data or information. For example, the acquisition unit 111 acquires conversational voice data 200 related to a conversation including utterances by multiple speakers. The conversational voice data 200 is data in which changes in electrical signals related to conversational voices are recorded in time series. The acquisition unit 111 transmits the acquired conversational voice data 200 to the estimation unit 112.
[0014] The estimation unit 112 estimates various data or information. For example, the estimation unit 112 estimates the start time and end time of an utterance by each dominant speaker based on the conversational voice data 200 transmitted from the acquisition unit 111. The estimation unit 112 transmits the estimated start time and end time to the identification unit 113 and the evaluation unit 114.
[0015] The identification unit 113 identifies various data or information. For example, the identification unit 113 identifies the timing of each dominant speaker's change (hereinafter also referred to as "speaker change") based on the start time and end time transmitted from the estimation unit 112. The identification unit 113 transmits the identified change timing to the evaluation unit 114.
[0016] The evaluation unit 114 evaluates various data or information. For example, the evaluation unit 114 evaluates the state of a conversation based on dialogue information D before and after a changeover timing transmitted from the identification unit 113. The dialogue information D includes at least one of the lengths of overlapping sections where speech sections of each dominant speaker overlap and the lengths of silent sections where speech sections of each dominant speaker do not overlap, as well as the lengths of the speech sections of each dominant speaker. The dialogue information D may include the start time and end time transmitted from the estimation unit 112. The evaluation unit 114 inputs the dialogue information D to a pre-trained detection model 120, thereby acquiring an evaluation result 300 regarding the state of the conversation from the detection model 120. The evaluation unit 114 outputs the acquired evaluation result 300.
[0017] 2 is a block diagram showing an example of the hardware configuration of the conversation evaluation device 1 according to this embodiment. For example, the conversation evaluation device 1 is a computer (e.g., a personal computer, a tablet terminal, or a smartphone). The conversation evaluation device 1 includes, as its components, a processing circuit 11, a memory circuit 12, an input IF 13, an output IF 14, and a communication IF 15. Each component is communicably connected to one another via a bus (BUS), which is a common signal communication path.
[0018] The processing circuit 11 is a circuit that controls the overall operation of the conversation evaluation device 1. The processing circuit 11 includes at least one processor. The processor refers to circuits such as a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), or a field programmable gate array (FPGA)). When the processor is a CPU, the CPU realizes each function by reading and executing each program stored in the storage circuit 12. When the processor is an ASIC, each function is directly incorporated into the ASIC as a logic circuit. The processor may be configured as a single circuit or may be configured by combining multiple independent circuits. The processing circuit 11 realizes an acquisition unit 111, an estimation unit 112, an identification unit 113, an evaluation unit 114, and a system control unit 115. The processing circuit 11 is an example of a processing unit.
[0019] The system control unit 115 controls various operations performed by the processing circuit 11. For example, the system control unit 115 provides an operating system (OS) for the processing circuit 11 to realize each unit (the acquisition unit 111, the estimation unit 112, the identification unit 113, and the evaluation unit 114).
[0020] The memory circuitry 12 is a circuit that stores various data or information. The memory circuitry 12 may be a processor-readable storage medium (e.g., a magnetic storage medium, an electromagnetic storage medium, an optical storage medium, or a semiconductor memory), or may be a drive that reads and writes data or information from and to the storage medium. The memory circuitry 12 stores each program that causes the processing circuitry 11 to realize each unit (the acquisition unit 111, the estimation unit 112, the identification unit 113, the evaluation unit 114, and the system control unit 115). The memory circuitry 12 may store a detection model 120. The memory circuitry 12 is an example of a memory unit.
[0021] The input IF 13 is an interface that accepts various inputs from a user. The input IF 13 converts the accepted inputs into electrical signals and transmits the electrical signals to the processing circuit 11. The input IF 13 may be a mouse, a keyboard, a button, a panel switch, a slider switch, a trackball, an operation panel, or a touch screen. The input IF 13 may be installed outside the conversation evaluation device 1. The input IF 13 is an example of an input unit.
[0022] The output IF14 is an interface that outputs various data or information to the user. The output IF14 outputs various data or information in response to an electrical signal transmitted from the processing circuit 11. The output IF14 may be a display device (e.g., a monitor) or an audio device (e.g., a speaker). The output IF14 may be installed outside the conversation evaluation device 1. The output IF14 is an example of an output unit, a display unit, or an audio unit.
[0023] The communication IF 15 is an interface for communicating various data or information with an external device, and is an example of a communication unit.
[0024] 3 is a flowchart showing an example of the overall operation of the conversation evaluation device 1 according to this embodiment. According to this example of operation, the conversation evaluation device 1 analyzes conversational voice data 200 and outputs an evaluation result 300. In particular, the conversation evaluation device 1 may start this example of operation in response to an instruction input by the user via the input IF 13. The conversation evaluation device 1 may acquire the conversational voice data 200 from the communication IF 15 and present the evaluation result 300 to the user via the output IF 14.
[0025] (Step S1) First, the conversation evaluation device 1 acquires conversational voice data 200 using the acquisition unit 111. The conversational voice data 200 is acquired by at least one sound collection device. For example, the conversational voice data 200 is voice data related to a telephone conference, a web conference, or a video conference held by multiple speakers using their own sound collection devices. The sound collection device may be a built-in microphone mounted on a handset, smartphone, or personal computer of each speaker. Alternatively, the sound collection device may be a headset microphone or a desktop microphone connected to each speaker's personal computer.
[0026] (Step S2) Next, the conversation evaluation device 1 determines, via the acquisition unit 111, whether or not the conversation voice data 200 acquired in step S1 has been source-separated. Specifically, the acquisition unit 111 determines whether or not the conversation voice data 200 has been separated into voices (separated voices) for each speaker. If the conversation voice data 200 has been source-separated (step S2-YES), the process proceeds to step S4. If the conversation voice data 200 has not been source-separated (step S2-NO), the process proceeds to step S3.
[0027] First, assume that the conversational voice data 200 is collected from each of the sound collection devices of multiple speakers. In this case, since each sound collection device collects voice data from a single speaker, the conversational voice data 200 contains the voice of each speaker individually. Therefore, the conversation evaluation device 1 determines that the conversational voice data 200 has been separated into sound sources.
[0028] Second, assume that the conversational voice data 200 is collected from a single sound collection device. In this case, the sound collection device simultaneously collects voices from multiple speakers, and therefore the conversational voice data 200 includes a voice (mixed voice) in which the voices of the speakers are mixed. Therefore, the conversation evaluation device 1 determines that the conversational voice data 200 has not been separated into sound sources.
[0029] (Step S3) Next, the conversation evaluation device 1 uses the acquisition unit 111 to separate the conversational voice data 200 that was determined not to have been source-separated in step S2. As a result, the conversational voice data 200 contains the voices of each speaker individually. The sound source separation can be achieved by known techniques (see Non-Patent Document 1).
[0030] (Step S4) Next, the conversation evaluation device 1 uses the estimation unit 112 to estimate the speech section of each speaker based on the voice of each speaker individually included in the conversational voice data 200. Specifically, the estimation unit 112 estimates the start time and end time of an utterance by each speaker and the length of the utterance section. The length of the utterance section is the duration from the start time to the end time of the utterance. The estimation of the speech section can be achieved by known technology (see Non-Patent Document 2).
[0031] (Step S5) Subsequently, the conversation evaluation device 1 estimates the utterances of each main speaker by the estimation unit 112 based on the speech sections of each speaker estimated in step S4 (see FIG. 4).
[0032] (Step S6) Next, the conversation evaluation device 1 uses the identification unit 113 to identify speaker changes for each dominant speaker based on the utterances of each dominant speaker estimated in step S5. Specifically, the identification unit 113 determines whether the two dominant speakers are different from each other for the speech periods of two temporally adjacent dominant speakers. If the two dominant speakers are different from each other, the identification unit 113 identifies that a speaker change has occurred between the speech periods of the two dominant speakers (see FIG. 5).
[0033] (Step S7) Finally, the conversation evaluation device 1 outputs, via the evaluation unit 114, an evaluation result 300 regarding the state of the conversation in the conversational voice data 200 for the speaker turn identified in step S6, based on the dialogue information D before and after the speaker turn (see FIG. 6). After step S7, the conversation evaluation device 1 ends the series of operations.
[0034] 4 is a flowchart showing a method for estimating a dominant speaker's utterance according to this embodiment. According to this estimation method, the estimation unit 112 estimates, for each speech section to be processed, whether the utterance is a dominant speaker's utterance or a non-dominant speaker's utterance.
[0035] (Step S51) First, the estimation unit 112 selects a speech section to be processed from all the speech sections estimated in step S4. For example, the estimation unit 112 selects the speech section with the earliest speech start time from all the unprocessed speech sections. That is, each time step S51 is executed, one speech section to be processed is selected.
[0036] (Step S52) Next, the estimation unit 112 determines whether the length of the utterance section selected in step S51 is equal to or greater than a threshold. This threshold is empirically determined (e.g., 0.5 seconds). If the length of the utterance section is equal to or greater than the threshold (step S52-YES), the process proceeds to step S53. If the length of the utterance section is not equal to or greater than the threshold (step S52-NO), the process proceeds to step S54A. That is, the estimation unit 112 estimates that the utterance section whose length is not equal to or greater than the threshold is a short utterance such as a backchannel, and is an utterance of a non-primary speaker.
[0037] (Step S53) Next, for an utterance section whose length is determined to be equal to or greater than the threshold in step S52, the estimation unit 112 determines whether the utterance section is included in another utterance section. Specifically, the estimation unit 112 determines whether the start time and end time of the utterance section are included between the start time and end time of the other utterance section. If the utterance section is included in another utterance section (step S53-YES), the process proceeds to step S54A. If the utterance section is not included in another utterance section (step S53-NO), the process proceeds to step S54B.
[0038] (Step S54A) In this case, the estimation unit 112 estimates that the speech section selected in step S51 is an utterance by a “non-dominant speaker.” After step S54A, the process proceeds to step S55.
[0039] (Step S54B) In this case, the estimation unit 112 estimates that the speech section selected in step S51 is an utterance by the “dominant speaker.” After step S54B, the process proceeds to step S55.
[0040] (Step S55) Subsequently, the estimation unit 112 determines whether all of the speech sections estimated in step S4 have been processed from step S51 to step S54A or S54B. If all of the speech sections have been processed (step S55-YES), the process proceeds to step S6 (see FIG. 3). If all of the speech sections have not been processed (step S55-NO), the process returns to step S51.
[0041] Note that the method for estimating the speech of a dominant speaker is not limited to the above method. The estimation unit 112 may recognize the content of each utterance using an external speech recognition system. When the recognized content of an utterance includes only a backchannel, the estimation unit 112 estimates that the utterance is the speech of a "non-dominant speaker." Conversely, the estimation unit 112 estimates that an utterance that is not estimated to be the speech of a non-dominant speaker is the speech of a "dominant speaker."
[0042] 5A and 5B are diagrams showing an example of conversation analysis according to this embodiment. Table 500A in Fig. 5A shows the analysis results of the conversation in the conversational audio data 200. Graph 500B in Fig. 5B graphically shows the analysis results of the conversation in table 500A. Table 500A or graph 500B may be displayed on the output IF 14 as a display device.
[0043] Table 500A shows a conversation between three speakers (A, B, and C). Table 500A shows utterances from each speaker along the rows, and items related to each utterance along the columns. The items are: column 1 "Dominant Speaker," column 2 "Speaker Change," column 3 "Start Time," column 4 "End Time," column 5 "Speaker," and column 6 "Content of Utterance."
[0044] The first column, "Dominant Speaker," indicates with a check mark the utterances of the dominant speaker estimated in step S5 of FIG. 3 (see FIG. 4). Specifically, of the 11 utterances included in table 500A, seven utterances, namely, the first, fourth to seventh, ninth, and eleventh utterances from the top, are estimated as utterances of the "dominant speaker." Conversely, four utterances, namely, the second to third, eighth, and tenth utterances from the top, are estimated as utterances of the "non-dominant speaker." In other words, utterances that include backchannels ("yes") or fragmentary repetitions ("chocolate") as utterance content are estimated as utterances of the "non-dominant speaker."
[0045] The second column, "Speaker Changes," indicates speaker changes of the dominant speaker identified in step S6 of Fig. 3 with stars 51, 52, 53, and 54. Specifically, of the 11 utterances included in table 500A, four utterances, the fifth, sixth, ninth, and eleventh from the top, are marked with stars 51, 52, 53, and 54. The stars 51, 52, 53, and 54 are marked on utterances after a speaker change.
[0046] Graph 500B shows three speakers (A, B, and C) along the vertical axis, and the speech periods of each speaker along the horizontal axis. Within each speech period, utterances by the "dominant speaker" are indicated by diagonally shaded bars 510, and utterances by the "minor speaker" are indicated by white bars 520. Each of the bars 510 and 520 corresponds to one of the 11 utterances included in table 500A. The four bars 510 marked with stars 51, 52, 53, and 54 in graph 500B correspond to the four utterances marked with stars 51, 52, 53, and 54 in table 500A.
[0047] 6 is a flowchart showing a method for evaluating the state of a conversation according to this embodiment. According to this evaluation method, the evaluation unit 114 outputs an evaluation result 300 based on dialogue information D before and after a speaker change.
[0048] (Step S71) First, for each speaker turn identified in step S6, the evaluation unit 114 extracts dialogue information D before and after the speaker turn.
[0049] For example, let us focus on the first speaker change marked with star 51 in FIG. 5. As can be seen from table 500A and graph 500B, this speaker change is from dominant speaker A to dominant speaker B. The length of the speech interval in dominant speaker A's turn before the speaker change is the length of the speech intervals of the first and fourth utterances from the top of table 500A. This length is calculated as (6.976-5.672)+(8.568-7.408)=2.464 (seconds). On the other hand, the length of the speech interval in dominant speaker B's turn after the speaker change is the length of the speech interval of the fifth utterance from the top of table 500A. This length is calculated as (9.576-8.728)=0.848 (seconds).
[0050] Furthermore, in the first speaker change marked with star 51, the speech period of dominant speaker A and the speech period of dominant speaker B do not overlap with each other. Specifically, according to the end time "8:568" of the fourth speech period from the top of table 500A and the start time "8:728" of the fifth speech period from the top, the two speech periods do not overlap with each other. Therefore, the presence or absence of an overlapping period is "no," and the length of the overlapping period is "0.000" (seconds). Conversely, there is a silent period between the two speech periods. Therefore, the presence or absence of a silent period is "yes," and the length of the silent period is calculated to be (8.728-8.568) = 0.160 (seconds).
[0051] Next, let us focus on the second speaker change marked with star 52 in FIG. 5. As can be seen from table 500A and graph 500B, this speaker change is from dominant speaker B to dominant speaker C. The length of the speech interval in dominant speaker B's turn before the speaker change is the length of the speech interval of the fifth utterance from the top in table 500A. This length is calculated as (9.576-8.728) = 0.848 (seconds). On the other hand, the length of the speech interval in dominant speaker C's turn after the speaker change is the length of the speech interval of the sixth and seventh utterances from the top in table 500A. This length is calculated as (9.757-8.800) + (11.829-10.821) = 1.965 (seconds).
[0052] Furthermore, in the second speaker change marked with star 52, the speech period of dominant speaker B and the speech period of dominant speaker C overlap with each other. Specifically, according to the end time "9:576" of the fifth speech period from the top of table 500A and the start time "8:800" of the sixth speech period from the top, the two speech periods overlap with each other. Therefore, the presence or absence of an overlapping period is "yes," and the length of the overlapping period is (9:576-8:800) = 0.776 (seconds). Conversely, there is no silent period between the two speech periods. Therefore, the presence or absence of a silent period is "no," and the length of the silent period is "0.000" (seconds).
[0053] Similarly, the evaluation unit 114 extracts dialogue information D for each of the turns marked with stars 53 and 54 (see FIG. 7).
[0054] (Step S72) Next, the evaluation unit 114 applies the trained detection model 120 to the dialogue information D for each speaker turn extracted in step S71. Specifically, the evaluation unit 114 inputs the dialogue information D for each speaker turn into the trained detection model 120, thereby detecting a predetermined event or probability for each speaker turn as the evaluation result 300.
[0055] The detection model 120 is trained with training data in which dialogue information D at a speaker change is used as input data and labels indicating the conversation state before and after the speaker change are used as ground truth data. For example, the detection model 120 is trained with training data in which dialogue information D is used as input data and the probability of a predetermined event occurring when this dialogue information D is given is used as ground truth data.
[0056] First, the detection model 120 is trained using paired data of the dialogue information D and a label indicating whether an authoritative speaker change has occurred. In this case, the trained detection model 120 can detect the "probability of an authoritative speaker change occurring." Second, the detection model 120 is trained using paired data of the dialogue information D and a label indicating the name of the dominant speaker who is making an authoritative utterance. In this case, the trained detection model 120 can detect the "probability of each dominant speaker making an authoritative utterance." Third, the detection model 120 is trained using paired data of the dialogue information D and a label indicating a value related to whether the conversation is active or not. In this case, the trained detection model 120 can detect the "level of conversational activity."
[0057] For example, let us consider the four speaker turns marked with stars 51, 52, 53, and 54 in Figure 5. Here, let us denote the dialogue information D in the Nth speaker turn as X n Let Y be the probability that a given event occurs in the Nth speaker turn. nAssume that (n: a natural number from 1 to N) the dialogue information D includes (1) the length of the speech interval of the dominant speaker before the speaker change, (2) the length of the speech interval of the dominant speaker after the speaker change, (3) the length of the silent interval, and (4) the length of the overlap interval. In this case, the dialogue information X1 for the first speaker change is expressed as X1 = {2.464, 0.848, 0.160, 0.000}. Similarly, the dialogue information X2, X3, and X4 for the second, third, and fourth speaker changes are expressed as X2 = {0.848, 1.965, 0.000, 0.776}, X3 = {1.965, 1.269, 0.851, 0.000}, and X4 = {1.269, 2.259, 1.176, 0.000}.
[0058] The evaluation unit 114 inputs dialogue information X1 to the trained detection model 120, which then outputs a probability Y1 that a predetermined event occurred in the first speaker turn. Similarly, the evaluation unit 114 inputs dialogue information X2, X3, and X4 to the trained detection model 120, which then outputs probabilities Y2, Y3, and Y4 that a predetermined event occurred in the second, third, and fourth speaker turns.
[0059] The detection model 120 may be a machine learning model (e.g., a regression model, a support vector machine, a decision tree, or a neural network). n The probability that a given event occurs for Y n Alternatively, the probability that a predetermined event has occurred may be output for each of a plurality of consecutive pieces of dialogue information.
[0060] Furthermore, the input data input to the detection model 120 may include features other than the dialogue information D. Specifically, the input data may include acoustic features (e.g., mean, variance) related to the pitch or power of the voice of the dominant speaker before and after the speaker change. Alternatively, the input data may include text information indicating the content of the utterance of each dominant speaker obtained by speech recognition.
[0061] (Step S73) Finally, the evaluation unit 114 evaluates the dialogue information X n A given event (with probability Y n ) to generate an evaluation result 300. For example, the evaluation unit 114 generates an evaluation result 300 based on the probability Y n If is greater than or equal to the threshold, the dialogue information X n The evaluation unit 114 generates alert information by associating it with the dialogue information X n The alert information generated for each evaluation is integrated to output the evaluation result 300. For example, the evaluation result 300 is displayed on the output IF 14 as a display device (see FIG. 8).
[0062] 7 is a diagram showing an example of dialogue information D according to this embodiment. Table 700 in Fig. 7 shows dialogue information D extracted for each of the four speaker turns in table 500A and graph 500B in Fig. 5.
[0063] Table 700 shows speaker changes marked with stars 51, 52, 53, and 54 along the rows, and items related to each speaker change along the columns. The items are: "Speaker change" in the first column, "Length of speech interval of dominant speaker before speaker change" in the second column, "Length of speech interval of dominant speaker after speaker change" in the third column, "Whether or not there was a silent interval" in the fourth column, "Length of silent interval" in the fifth column, "Whether or not there was an overlapping interval" in the sixth column, and "Length of overlapping interval" in the seventh column.
[0064] 8 is a diagram showing a display example of the evaluation result 300 according to this embodiment. After the telephone conference, web conference, or video conference has ended, the conversation evaluation device 1 may perform the operation of FIG. 3 when the audio recording of the conversation has finished, thereby outputting the evaluation result 300 afterwards (offline operation). Alternatively, the conversation evaluation device 1 may perform the operation of FIG. 3 while acquiring the conversation audio data 200 in real time during the telephone conference, web conference, or video conference, thereby outputting the evaluation result 300 in real time (real-time operation).
[0065] (Example of Offline Operation) For example, assume that three speakers (A, B, and C) are holding a conference in a conference room. Furthermore, assume that during the conference, speaker B starts speaking as if interrupting speaker A, and then speaker B continues to speak unilaterally for a long time. After the conference ends, the conversation evaluation device 1 performs the operation shown in FIG. 3 based on conversational voice data 200 collected by a sound collection device installed in the conference room.
[0066] Graph 800 shows three speakers (A, B, C) along the vertical axis, and the speech intervals of each speaker along the horizontal axis. Each speech interval is indicated by a diagonally shaded bar 81. In particular, around time "05m00s", a speaker change occurs from speaker A to speaker B. If an intimidating speaker change is detected based on dialogue information D before and after this speaker change, alert information is displayed in association with the speaker change. For example, the alert information is represented by a box 82.
[0067] (Example 1 of Real-Time Operation) Similarly, assume that three speakers (A, B, C) are holding a conference in a conference room. The conversation evaluation device 1 performs the operation of FIG. 3 based on conversational audio data 200 collected in real time by a sound collection device installed in the conference room. If the conversation evaluation device 1 detects intimidating speaker changes more than a threshold number of times during the conference, it outputs alert information indicating that "intimidating speaker changes may have occurred during the conference." For example, the conversation evaluation device 1 sends an email containing this alert information to a terminal of the supervisor or human resources department who manages the three speakers (A, B, C).
[0068] (Example 2 of Real-Time Operation) For example, assume that three speakers (A, B, and C) are holding an online conference in the presence of one facilitator F. The conversation evaluation device 1 performs the operation of FIG. 3 based on conversational voice data 200 collected in real time during the conference. The conversation evaluation device 1 transmits an evaluation result 300 indicating the activity level of the conference to the terminal of facilitator F in real time. Furthermore, if the activity level of the conference is below a threshold for a predetermined time or number of times, the conversation evaluation device 1 outputs alert information indicating that "the conference is inactive." For example, the conversation evaluation device 1 sends an email containing this alert information to the terminal of facilitator F.
[0069] Fig. 9 shows an example of the estimation accuracy of conversation analysis using the conventional method and the proposed method. Table 900 shows the estimation accuracy of the conversation activity level of three people using two detection models 120 trained using different training methods. XGboost (eXtreme Gradient Boosting), a decision tree method, was used for the detection model 120.
[0070] To compare the conventional and proposed methods, a dataset of 10 sessions of casual conversations between three speakers was prepared. This dataset contained a total of 6,280 speaker turns, each of which was manually labeled as either "active" or "inactive."
[0071] The conventional method recorded (1) the length of the overlapping section as dialogue information D for each of the above speaker changes. The proposed method recorded (1) the length of the overlapping section, (2) the length of the silent section, (3) the length of the speech section of the dominant speaker before the speaker change, and (4) the length of the speech section of the dominant speaker after the speaker change as dialogue information D for each of the above speaker changes. As a result, paired data containing dialogue information D and labels was created for each of the 6,280 speaker changes of the dominant speaker.
[0072] Next, 80% of all paired data was divided as training data, and the remaining 20% was divided as evaluation data. A detection model 120 was trained using this training data. The trained detection model 120 estimated the label corresponding to each piece of dialogue information D contained in the evaluation data as "active" or "inactive." The estimated label was compared with the correct label to calculate the accuracy of the label estimation by the trained detection model 120. The F-measure was used as an evaluation scale for the accuracy of the label estimation. The closer the F-measure is to "1.0," the higher the accuracy of the label estimation.
[0073] According to table 900, the F-value for the label "active" using the detection model 120 trained using the conventional method is "0.46," and the F-value for the label "inactive" is "0.76." On the other hand, the F-value for the label "active" using the detection model 120 trained using the proposed method is "0.66," and the F-value for the label "inactive" is "0.88." In other words, it can be seen that the detection model 120 trained using the proposed method has higher estimation accuracy for the labels "active" and "inactive" compared to the conventional method.
[0074] According to the present embodiment described above, the conversation evaluation device 1 can appropriately evaluate the state of communication between a plurality of speakers. In particular, the conversation evaluation device 1 can objectively evaluate the state of conversation between a plurality of speakers.
[0075] First, when two speech sections overlap during a speaker change by a dominant speaker, the conversation evaluation device 1 uses the length of the speech section of the first speaker before the overlap occurs and the length of the speech section of the second speaker after the overlap occurs. This allows the conversation evaluation device 1 to evaluate which of the first and second speakers is speaking unilaterally. In other words, the conversation evaluation device 1 can evaluate which of the first and second speakers is speaking in an intimidating manner.
[0076] Second, the conversation evaluation device 1 can evaluate the activity level of the conversation between two speakers by using the lengths of the speech intervals of the two speakers and the lengths of the silent intervals when the two speakers take turns speaking. For example, if the silent intervals are short and the lengths of the speech intervals of the two speakers are long, the conversation evaluation device 1 can evaluate that the conversation between the two speakers is active. Conversely, if the silent intervals are long and the lengths of the speech intervals of the two speakers are short, the conversation evaluation device 1 can evaluate that the conversation between the two speakers is inactive.
[0077] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0078] 1...conversation evaluation device, 11...processing circuit, 12...memory circuit, 13...input IF, 14...output IF, 15...communication IF, 51, 52, 53, 54...stars, 81, 510, 520...bars, 82...box, 111...acquisition unit, 112...estimation unit, 113...identification unit, 114...evaluation unit, 115...system control unit, 120...detection model, 200...conversational voice data, 300...evaluation results, 500A, 700, 900...table, 500B, 800...graph
Claims
1. On the computer, estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation function for estimating the start time and end time of each utterance by a main speaker, by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold value and which is not included in another utterance period, as an utterance by a main speaker; and an identification function for identifying timings for the respective dominant speakers to take turns based on the estimated start and end times; an evaluation function that uses dialogue information before and after the specified takeover timing as input data, and a trained model that has been trained with training data in which labels indicating whether an intimidating takeover occurred before and after the takeover timing are used as correct answer data, to evaluate the probability that an intimidating takeover occurred as an evaluation result of the state of the conversation; and To achieve this, the dialogue information includes at least one of a length of an overlapping section in which the speech sections of the respective main speakers overlap and a length of a silent section in which the speech sections of the respective main speakers do not overlap, and a length of the speech section of each of the main speakers. Conversation assessment program.
2. On the computer, estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation function for estimating the start time and end time of each utterance by a main speaker, by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold value and which is not included in another utterance period, as an utterance by a main speaker; and an identification function for identifying timings for the respective dominant speakers to take turns based on the estimated start and end times; an evaluation function that uses dialogue information before and after the specified changeover timing as input data, and uses a trained model trained with training data in which labels indicating the names of dominant speakers who are making intimidating utterances before and after the changeover timing are used as correct answer data to evaluate the probability that each dominant speaker is making intimidating utterances as an evaluation result regarding the state of the conversation; To achieve this, the dialogue information includes at least one of a length of an overlapping section in which the speech sections of the respective main speakers overlap and a length of a silent section in which the speech sections of the respective main speakers do not overlap, and a length of the speech section of each of the main speakers. Conversation assessment program.
3. On the computer, estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation function for estimating the start time and end time of each utterance by a main speaker, by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold value and which is not included in another utterance period, as an utterance by a main speaker; and an identification function for identifying timings for the respective dominant speakers to take turns based on the estimated start and end times; an evaluation function that evaluates the conversation activity level as an evaluation result regarding the state of the conversation using a trained model that has been trained with training data in which dialogue information before and after the specified turn-taking timing is used as input data and labels indicating whether the conversation is active before and after the turn-taking timing are used as correct answer data; and To achieve this, the dialogue information includes at least one of a length of an overlapping section in which the speech sections of the respective main speakers overlap and a length of a silent section in which the speech sections of the respective main speakers do not overlap, and a length of the speech section of each of the main speakers. Conversation assessment program.
4. On the computer, estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation function for estimating the start time and end time of each utterance by a main speaker, by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold value and which is not included in another utterance period, as an utterance by a main speaker; and an identification function for identifying timings for the respective dominant speakers to take turns based on the estimated start and end times; an evaluation function that, when two speech sections overlap at the specified changeover timing, evaluates which of the first speaker and the second speaker is speaking in an intimidating manner as an evaluation result of the state of the conversation, using the length of the speech section of the first speaker before the two speech sections overlap and the length of the speech section of the second speaker after the two speech sections overlap; A conversation assessment program that makes this possible.
5. On the computer, estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation function for estimating the start time and end time of each utterance by a main speaker, by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold value and which is not included in another utterance period, as an utterance by a main speaker; and an identification function for identifying timings for the respective dominant speakers to take turns based on the estimated start and end times; an evaluation function that evaluates the conversation activity of the two speakers as an evaluation result regarding the state of the conversation, using the length of the speech intervals of the two speakers at the specified changeover timing and the length of the silent interval where the speech intervals of the two speakers do not overlap; A conversation assessment program that makes this possible.
6. 4. The conversation evaluation program according to claim 1, wherein, when an intimidating speaker change is detected based on the dialogue information, the evaluation function outputs alert information in association with the timing of the change at which the intimidating speaker change was detected.
7. the evaluation function outputs alert information when it detects an intimidating speaker change in the conversation a number of times equal to or greater than a threshold based on the dialogue information. The conversation evaluation program according to any one of claims 1 to 3.
8. the evaluation function outputs alert information when it detects that the conversation has been inactive for a number of times equal to or greater than a threshold based on the dialogue information. The conversation evaluation program according to any one of claims 1 to 3.
9. estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; an estimation unit that estimates start times and end times of utterances by each of the speakers by determining, among the utterance periods of each speaker, an utterance period whose length is equal to or greater than a threshold and which is not included in another utterance period as an utterance by a dominant speaker; an identification unit that identifies timings for changing the dominant speaker based on the estimated start time and end time; an evaluation unit that uses dialogue information before and after the specified take-over timing as input data, and uses a trained model that has been trained with training data in which labels indicating whether an intimidating take-over occurred before and after the take-over timing are used as correct answer data to evaluate the probability that an intimidating take-over occurred as an evaluation result regarding the state of the conversation; Equipped with the dialogue information includes at least one of a length of an overlapping section in which the speech sections of the respective main speakers overlap and a length of a silent section in which the speech sections of the respective main speakers do not overlap, and a length of the speech section of each of the main speakers. Conversation assessment device.
10. A computer comprising: estimating speech sections of each speaker based on speech data relating to a conversation including utterances by a plurality of speakers; determining, among the speech periods of each speaker, a speech period whose length is equal to or greater than a threshold and which is not included in another speech period as a speech period by a dominant speaker, and estimating a start time and an end time of each speech period by the dominant speaker; determining a timing for each of the dominant speakers to take over based on the estimated start and end times; and Using a trained model trained with training data in which dialogue information before and after the specified take-over timing is used as input data and labels indicating whether an intimidating take-over occurred before and after the take-over timing is used as correct answer data, evaluate the probability that an intimidating take-over occurred as an evaluation result of the state of the conversation; Equipped with the dialogue information includes at least one of a length of an overlapping section in which the speech sections of the respective main speakers overlap and a length of a silent section in which the speech sections of the respective main speakers do not overlap, and a length of the speech section of each of the main speakers. Conversation assessment methods.
Citation Information
Patent Citations
Information accumulation device
JP1998191245A
Voice analysis system and voice analysis device
JP2013072979A
Speech processor, speech processing method and speech processing program
JP2016133774A
Harassment prevention system and harassment prevention method
JP2023009563A
Audio processing device, audio processing method, and audio processing program
JP6524674B2