Speech evaluation system, speech evaluation method, and computer recording medium

By using a wearable terminal to detect speech and audience response in multiple participants' communication, the problem of insufficient detection accuracy of speech evaluation in the prior art is solved, and the accuracy of evaluation value calculation for each speech is improved.

CN114550723BActive Publication Date: 2025-06-20TOYOTA JIDOSHA KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111365493.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-11-19
Filing Date
2021-11-18
Publication Date
2025-06-20
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy in the evaluation and detection of speeches in exchanges between multiple participants, making it difficult to accurately reflect the audience's response to each speech.

Method used

A speech evaluation system is designed, and the speech evaluation value is calculated by wearing it on the participants through multiple wearable terminals. The sound collecting unit and acceleration sensor are used to detect the speech and the audience's reactions. The system not only considers the audience response during speech, but also includes the response to speech delay, and calculates the evaluation value by setting the start and end timing of the evaluation period.

Benefits of technology

The evaluation value calculation accuracy of each speech is improved, which can accurately reflect the audience's overall response to each speech, including timely and delayed responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550723B_ABST
    Figure CN114550723B_ABST
Patent Text Reader

Abstract

The present invention relates to a speech evaluation system, a speech evaluation method, and a computer-readable medium storing a program. A speech detection unit detects a speech during a conversation based on output values of microphones of a plurality of wearable terminals, and determines a wearable terminal corresponding to the detected speech. A speech period detection unit detects a start timing and an end timing of each speech detected by the speech detection unit. An evaluation value calculation unit calculates an evaluation value for each speech detected by the speech detection unit based on output values of acceleration sensors of wearable terminals other than the wearable terminal corresponding to the speech during a speech evaluation target period, where the speech evaluation target period is a period from a first timing after the start timing of the speech and before the end timing to a second timing after the end timing of the speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a speech evaluation system, a speech evaluation method, and a program. Background Art

[0002] In a communication composed of multiple participants, there is a need to extract important speeches that the audience particularly agrees with among the speeches of each of the multiple participants.

[0003] As such a technique, Patent Document 1 (Japanese Patent Application Laid-Open No. 2016-103081) counts the number of times the audience nods during the speech of a specific speaker using a wearable terminal worn by the specific speaker in a conversation among multiple users, and calculates the audience acceptance for the specific speaker based on the value obtained by dividing the number of times the audience nods by the speech time of the specific speaker (paragraphs 0080, 0093). Then, the higher the audience acceptance, the more the speech is accepted by the audience. Summary of the Invention

[0004] In the technique of Patent Document 1, there is room for improvement in the detection accuracy of the evaluation of the speech.

[0005] An object of the present invention is to provide a technique for accurately obtaining an evaluation value for each speech in a communication composed of multiple participants.

[0006] According to a first aspect of the present application invention, there is provided a speech evaluation system that obtains an evaluation value for each speech in a communication composed of multiple participants, wherein the speech evaluation system includes: a plurality of wearable terminals, each worn by each of the multiple participants, and each having at least a sensor including a sound collection unit; a speech detection unit that detects a speech in the communication based on output values of the sound collection units of the plurality of wearable terminals and determines the wearable terminal corresponding to the detected speech; a speech period detection unit that, for each speech detected by the speech detection unit, detects a start timing and an end timing of the speech; and an evaluation value calculation unit that, for each speech detected by the speech detection unit, calculates an evaluation value for the speech based on output values of the sensors of the wearable terminals other than the wearable terminal corresponding to the speech during a speech evaluation target period, where the speech evaluation target period is a period from a first timing after the start timing of the speech and before the end timing to a second timing after the end timing of the speech. According to the above structure, in addition to the reaction of the audience during the speech, the reaction of the audience that occurs after the speech is also reflected in the calculation of the evaluation value for the speech, so that an evaluation value for each speech can be accurately obtained.

[0007] Preferably, the second timing is set to a timing after a predetermined time has elapsed since the end timing of the corresponding speech. With the above structure, the calculation required to set the second timing is simplified, so the second timing can be set at low cost.

[0008] Preferably, the second timing is set to the timing when another speech starts after the corresponding speech. With the above structure, the evaluation value can be calculated excluding the reaction to another speech, so the evaluation value for the corresponding speech can be obtained with good accuracy.

[0009] Preferably, the second timing is set to a timing after a predetermined time has elapsed since the end timing of the corresponding speech. When another speech starts after the corresponding speech before the predetermined time has elapsed since the end timing of the corresponding speech, the second timing is set to the timing when another speech starts after the corresponding speech. With the above structure, when another speech does not start after the corresponding speech before the predetermined time has elapsed since the end timing of the corresponding speech, the second timing can be set at low cost, and when another speech starts after the corresponding speech before the predetermined time has elapsed since the end timing of the corresponding speech, the evaluation value can be calculated excluding the reaction to another speech, so the evaluation value for the corresponding speech can be obtained with good accuracy.

[0010] Preferably, the sensor includes an acceleration sensor.

[0011] Preferably, when the output value of the acceleration sensor indicates an action in which the participant wearing the corresponding wearable terminal swings the head longitudinally, the evaluation value calculation unit calculates the evaluation value in a manner that increases the evaluation value for the corresponding speech.

[0012] Preferably, when the output value of the acceleration sensor indicates an action in which the participant wearing the corresponding wearable terminal swings the head laterally, the evaluation value calculation unit calculates the evaluation value in a manner that decreases the evaluation value for the corresponding speech.

[0013] According to a second aspect of the invention of the present application, there is provided a speech evaluation method. In a communication composed of multiple participants, an evaluation value for each speech is obtained. A plurality of wearable terminals are worn by each of the multiple participants, and each of the plurality of wearable terminals has a sensor including at least a sound collection unit. The speech evaluation method includes: detecting a speech in the communication according to the output value of the sound collection unit of the plurality of wearable terminals, and determining the wearable terminal corresponding to the detected speech; for each detected speech, detecting the start timing and the end timing of the speech; and for each detected speech, calculating an evaluation value for the speech according to the output value of the sensor of the wearable terminals other than the wearable terminal corresponding to the speech during a speech evaluation target period, where the speech evaluation target period is a period from a first timing after the start timing of the speech and earlier than the end timing to a second timing later than the end timing of the speech. According to the above method, in addition to the reactions of the audience during the speech, the reactions of the audience that occur after the speech is delayed are also reflected in the calculation of the evaluation value for the speech, so that an evaluation value for each speech can be accurately obtained.

[0014] In addition, there is provided a computer recording medium recording a program for causing a computer to execute the above speech evaluation method.

[0015] According to the present invention, in addition to the reactions of the audience during the speech, the reactions of the audience that occur after the speech is delayed are also reflected in the calculation of the evaluation value for the speech, so that an evaluation value for each speech can be accurately obtained.

[0016] The above and other objects, features, and advantages of the present disclosure will be more fully understood from the following detailed description and the accompanying drawings given by way of illustration only, and should not be considered as a limitation of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic diagram of a speech evaluation system.

[0018] Figure 2 is a functional block diagram of a wearable terminal.

[0019] Figure 3 is a diagram illustrating the structure of transmitted data.

[0020] Figure 4 is a functional block diagram of an evaluation device.

[0021] Figure 5 is a diagram showing the transmitted data stored in the evaluation device.

[0022] Figure 6It is a graph showing the detected speech and approval ratio.

[0023] Figure 7 It is a graph showing a monotonically increasing function for emphasis processing.

[0024] Figure 8 It is a graph showing the detected speech and the f(p) value.

[0025] Figure 9 It is a diagram illustrating the structure of the evaluation data.

[0026] Figure 10 It is the control flow of the speech evaluation system.

[0027] Figure 11 It is a graph showing a step function for emphasis processing. Detailed Implementation Modes

[0028] Hereinafter, the present invention will be described by way of embodiments of the invention, but the invention related to the patent claims is not limited to the following embodiments. In addition, not all of the structures described in the embodiments are necessarily required as technical means for solving the problems. For the sake of clarity of explanation, the following description and drawings have been appropriately omitted and simplified. In each drawing, the same reference numerals are assigned to the same elements, and repeated explanations are omitted as needed.

[0029] In Figure 1 a schematic diagram of the speech evaluation system 1 is shown. The speech evaluation system 1 is a system that obtains an evaluation value for each speech in a communication composed of a plurality of participants 2. The speech evaluation system 1 includes a plurality of wearable terminals 3 and an evaluation device 4.

[0030] In the present embodiment, the number of participants 2 constituting the same communication is set to 3, but it is not limited thereto, and it may be two or four or more, for example, it may be 10. Regarding the communication, typically it is a conversation-type communication established by the speech of each person. Such a communication is, for example, a seminar, a symposium, or a training session. However, the communication is not limited to the communication in which all participants gather in the same real space, and may also include the communication in which they gather in a virtual space online.

[0031] (Wearable Terminal 3)

[0032] As Figure 1As shown, multiple wearable terminals 3 are respectively worn by multiple participants 2 for use. That is, one participant 2 wears one wearable terminal 3. In the present embodiment, the wearable terminal 3 is a badge that can be attached to and detached from the upper garment worn on the upper body of the participant 2, and preferably is installed at a position above the chest (heart). However, the wearable terminal 3 may not be a badge, but a head-mounted earphone, earphone, glasses, necklace, pendant, etc.

[0033] Figure 2 A functional block diagram of each wearable terminal 3 is shown. As Figure 2 shown, the wearable terminal 3 includes a terminal ID information storage unit 10, a microphone 11, and an acceleration sensor 12. The wearable terminal 3 also includes a CPU 3a (Central Processing Unit) as a central arithmetic processor, a RAM 3b (Random Access Memory) that can be read and written freely, and a ROM 3c (Read Only Memory) that is for reading only. Moreover, the CPU 3a reads the control program stored in the ROM 3c and executes it, so that the control program causes the hardware such as the CPU 3a to function as a time counting unit 13, a transmission data generation unit 14, and a data transceiver unit 15. Each wearable terminal 3 can perform two-way wireless communication with the evaluation device 4 via the data transceiver unit 15.

[0034] The terminal ID information storage unit 10 stores terminal ID information for identifying the corresponding wearable terminal 3 from other wearable terminals 3. Regarding the terminal ID information, typically, the MAC address unique to the wearable terminal 3 can be cited. However, the terminal ID information may also be a number, character, or a combination thereof set by the evaluation device 4 at the startup of the wearable terminal 3. In the present embodiment, the terminal ID information is set to a natural number set by the evaluation device 4 at the startup of the wearable terminal 3.

[0035] The microphone 11 is a specific example of a sound collection unit, which converts the sound around the corresponding wearable terminal 3 into a voltage value and outputs it to the transmission data generation unit 14.

[0036] The acceleration sensor 12 converts the three-axis acceleration of the corresponding wearable terminal 3 into a voltage value and outputs it to the transmission data generation unit 14. When the participant 2 wearing the corresponding wearable terminal 3 swings their head "vertically", the upper body of the participant 2 repeats bending and stretching around the roll axis (an axis parallel to the axis connecting the left shoulder and the right shoulder). Therefore, in this case, the vertical component value in the output value of the acceleration sensor 12 varies in a manner of repeating increase and decrease within a predetermined range. On the other hand, when the participant 2 wearing the corresponding wearable terminal 3 swings their head "horizontally", the upper body of the participant 2 repeats twisting around the yaw axis (an axis parallel to the direction in which the spine extends). Therefore, in this case, the output value corresponding to the horizontal component value in the output value of the acceleration sensor 12 varies in a manner of repeating increase and decrease within a predetermined range.

[0037] The microphone 11 and the acceleration sensor 12 constitute the sensor 16 for detecting the words and deeds of the participant 2 wearing the corresponding wearable terminal 3. However, the acceleration sensor 12 may be omitted.

[0038] The time counting unit 13 has time data, increments the time data initialized by a predetermined method at a predetermined period, and outputs the time data to the transmission data generation unit 14. The time data possessed by the time counting unit 13 is typically initialized using the time data received from the evaluation device 4. Instead of this method, the time data possessed by the time counting unit 13 may be initialized by the corresponding wearable terminal 3 accessing the Network Time Protocol (NTP) via the evaluation device 4 and the Internet to obtain the latest time data.

[0039] The transmission data generation unit 14 generates Figure 3 the shown transmission data 14a at a predetermined interval. As Figure 3 shown, the transmission data 14a includes terminal ID information, time data, sound data, and acceleration data. The predetermined interval is typically 1 second. The sound data is the output value of the microphone 11 from the time indicated by the time data until 1 second has elapsed. Similarly, the acceleration data is the output value of the acceleration sensor 12 from the time indicated by the time data until 1 second has elapsed.

[0040] Return to Figure 2, the data transceiver unit 15 sends the transmission data 14a to the evaluation device 4. In the present embodiment, the data transceiver unit 15 sends the transmission data 14a to the evaluation device 4 through short-range wireless communication such as Bluetooth (registered trademark). However, instead of this method, the data transceiver unit 15 may send the transmission data 14a to the evaluation device 4 through wired communication. In addition, the data transceiver unit 15 may also send the transmission data 14a to the evaluation device 4 via a network such as the Internet.

[0041] (Evaluation device 4)

[0042] Figure 4 A functional block diagram of the evaluation device 4 is shown. As Figure 4 shown, the evaluation device 4 includes a CPU 4a (Central Processing Unit, central processing unit) as a central arithmetic processor, a RAM 4b (Random Access Memory, random access memory) that can be read and written freely, and a ROM 4c (Read Only Memory, read-only memory) for reading only. Then, the CPU 4a reads the control program stored in the ROM 4c and executes it, so that the control program causes the hardware such as the CPU 4a to function as a data transceiver unit 20, a data storage unit 21, a speech detection unit 22, a speech period detection unit 23, an approval ratio calculation unit 24, an emphasis processing unit 25, an evaluation value calculation unit 26, and an evaluation value output unit 27.

[0043] The data transceiver unit 20 receives the transmission data 14a from each wearable terminal 3 and stores the received transmission data 14a in the data storage unit 21. Figure 5 The transmission data 14a stored in the data storage unit 21 is shown. As Figure 5 shown, in the data storage unit 21, the transmission data 14a received from each wearable terminal 3 is directly stored in the order of reception.

[0044] Return to Figure 4 , the speech detection unit 22 detects the speech during the conversation based on the output values of the microphones 11 of the plurality of wearable terminals 3, and determines the wearable terminal 3 corresponding to the detected speech.

[0045] Specifically, the speech detection unit 22 analyzes the voice data stored in the data storage unit 21. When the voice data of any one of the plurality of transmission data 14a at a certain moment exceeds a predetermined value, it is detected that there is speech during the conversation at that certain moment, and with reference to the terminal ID information of the transmission data 14a, the wearable terminal 3 corresponding to the detected speech is determined.

[0046] Figure 6Examples of the speech detection unit 22 detecting speech a, speech b, speech c, and speech d. Figure 6 The horizontal axis of Figure 6 is time. The speech detection unit 22 detects speech a, speech b, speech c, and speech d in this recorded order without repetition. Speech a and speech c are speeches made by the participant 2 wearing the wearable terminal 3 with terminal ID: 1. Similarly, speech b is a speech made by the participant 2 wearing the wearable terminal 3 with terminal ID: 2, and speech d is a speech made by the participant 2 wearing the wearable terminal 3 with terminal ID: 3.

[0047] In addition, the method by which the speech detection unit 22 detects speech and determines the corresponding wearable terminal 3 is not limited to the above method.

[0048] For example, when the voice data of any one of the multiple transmission data 14a at a certain moment is greater than a predetermined amount than the voice data of the other transmission data 14a at that moment, it is detected that there is speech in the communication at that certain moment, and by referring to the terminal ID information of the transmission data 14a, the wearable terminal 3 corresponding to the detected speech can be determined.

[0049] In addition, as preprocessing for detecting speech, the speech detection unit 22 can also remove the steady-state noise included in the voice data. Steady-state noise refers to, for example, the operating sound of an air conditioner or noise caused by surrounding noise. In addition, as preprocessing for detecting speech, the speech detection unit 22 can also remove the non-steady-state noise included in the voice data. Non-steady-state noise refers to noise caused by sudden loud voices of non-participants who did not participate in the communication or object sounds caused by the opening and closing of doors. Such non-steady-state noise has the property of occurring at almost the same level as the voice data of the multiple transmission data 14a at a certain moment.

[0050] Return to Figure 4 , for each speech detected by the speech detection unit 22, the speech period detection unit 23 detects the start timing and end timing of the speech. In the example of Figure 6 , the start timing of speech a is time t1, and the end timing is time t2. The start timing of speech b is time t4, and the end timing is time t5. The start timing of speech c is time t6, and the end timing is time t7. The start timing of speech d is time t8, and the end timing is time t9. In addition, in this specification, "timing" is a concept that determines a certain time point on the time axis, and it can be either a time composed of hours, minutes, and seconds or a simple natural number that increases as time passes. Therefore, in this specification, "timing" can also be simply referred to as "time".

[0051] Return to Figure 4, the approval ratio calculation unit 24 calculates the approval ratio at every predetermined time interval. Here, the approval ratio is a ratio obtained by dividing the number of nodding listeners among multiple listeners by the total number of listeners, and is a value between 0 and 1. The predetermined time interval is, for example, 5 seconds. When this time interval is too large, nodding actions at different timings are treated as nodding actions at the same timing, so the approval actions for the speech are over-evaluated. When this time interval is too small, nodding actions at almost the same timing are treated as nodding actions at different timings, so the approval actions for the speech are under-evaluated.

[0052] The approval ratio calculation unit 24 first refers to Figure 5 the stored transmission data 14a shown, and calculates the approval ratio during the process of speaking a. That is, the approval ratio calculation unit 24 analyzes the acceleration data of the transmission data 14a corresponding to the terminal ID: 2 during the period from t1 to 5 seconds elapsed, and determines whether the participant 2 wearing the wearable terminal 3 corresponding to the terminal ID: 2 has made a nodding action. A specific example of determining the presence or absence of a nodding action based on the acceleration data is as follows.

[0053] That is, the approval ratio calculation unit 24 extracts the vertical component values of the acceleration data during the period from time t1 to 5 seconds elapsed, calculates the average value and standard deviation of the extracted vertical component values, and when the standard deviation is smaller than a predetermined value and there is a vertical component value deviating from the average value by a predetermined amount in a single-occurrence manner, it is determined that the participant 2 wearing the wearable terminal 3 corresponding to the terminal ID: 2 has made a nodding action during the period from time t1 to 5 seconds elapsed. The same applies to the terminal ID: 3. The approval ratio calculation unit 24 repeats the above calculation of the approval ratio in the same way after 5 seconds have elapsed since time t1, and ends at the time t4 when there is a speech other than speech a.

[0054] When the approval ratio calculation unit 24 determines the presence or absence of a nodding action, on the premise that the standard deviation of the vertical component values of the acceleration data is smaller than a predetermined value, it is possible to remove noise caused by large actions other than nodding actions such as the walking actions and posture changes of the participant 2.

[0055] In Figure 6 the example, from time t1 to time t2, the approval ratio once rose sharply from around 0, and after a temporary decline, it rose again. The approval ratio remained constant around time t2, and then, before reaching time t4, it roughly returned to zero.

[0056] In addition, in Figure 6In the example, the number of the audience is only two, so the approval ratio could originally be any value among 0, 0.5, and 1.0. However, for the sake of better understanding, the approval ratio is made to change slowly as if the number of the audience is around 30 people.

[0057] Next, the approval ratio calculation unit 24 calculates the approval ratio during the speech b. That is, the approval ratio calculation unit 24 analyzes the acceleration data of the transmission data 14a corresponding to the terminal ID: 1 from the time t4 until 5 seconds have elapsed, and determines whether the participant 2 wearing the wearable terminal 3 corresponding to the terminal ID: 1 has nodded. The same applies to the terminal ID: 3. The approval ratio calculation unit 24 also repeats the above calculation of the approval ratio after 5 seconds from the time t4, and ends at the time t6 when there is a speech other than the speech b.

[0058] In Figure 6 the example, from the time t4 to the time t5, the approval ratio changes with a value less than 0.5, and around the time t5, it is approximately 0.

[0059] The approval ratio calculation unit 24 also calculates the approval ratio after the time t6 in the same manner.

[0060] In addition, as another method for the approval ratio calculation unit 24 to determine whether there is a nodding motion, the vertical component value can be extracted from the transmission data 14a every predetermined time interval, and the extracted vertical component value is input into a learned Convolution Neural Network (CNN). When the output value of the convolutional neural network is equal to or greater than a predetermined value, it is determined that the participant 2 wearing the wearable terminal 3 has nodded during this time interval. Additionally, as another method for the approval ratio calculation unit 24 to determine whether there is a nodding motion, the vertical component value can be extracted from the transmission data 14a every predetermined time interval, and various characteristic quantities (such as the difference between the maximum value and the minimum value, the variance value, the frequency distribution, etc.) of the extracted vertical component value are calculated. The calculated characteristic quantities are input into a learned support vector machine (SVM), and its output value is used.

[0061] Returning to Figure 4 , the emphasis processing unit 25 performs emphasis processing on the approval ratio calculated by the approval ratio calculation unit 24 to emphasize the high or low of the approval ratio. In the emphasis processing, for example, the following formula (1) which is a monotonically increasing function can be used. Here, p represents the approval ratio, and k is an adjustment parameter.

[0062]

Formula 1

[0063]

[0064] Figure 7 It is the graph of the above formula (1) used in the emphasis process performed by the emphasis processing unit 25. The horizontal axis represents the approval ratio, and the vertical axis represents the f(p) value. The larger the adjustment parameter k is, the more sharply convex the curve depicted by the f(p) value becomes towards the lower right on the graph. According to the emphasis process based on the above formula (1), when the audience nods almost unanimously, the f(p) value becomes a large value, and when the audience nods sporadically at different timings, the f(p) value becomes a small value. Through such an emphasis process, important speeches such as those where the audience nods almost unanimously can be made more prominent than relatively unimportant speeches.

[0065] Figure 8 Shows the f(p) value after the emphasis process. According to Figure 8 , when the audience does not nod almost unanimously, even in the time interval when a certain degree of the audience nods, the f(p) value in that time interval is halved or compressed to a value close to zero.

[0066] The evaluation value calculation unit 26 sets an evaluation period as the speech evaluation object period corresponding to each speech detected by the speech detection unit 22, and calculates the evaluation value for that speech.

[0067] (Speech a)

[0068] Specifically, the evaluation value calculation unit 26 sets the start timing (the first timing) of the evaluation period corresponding to speech a after the start timing of speech a, i.e., time t1, and before the end timing, i.e., time t2. In this embodiment, the evaluation value calculation unit 26 sets the start timing of the evaluation period corresponding to speech a as the start timing of speech a, i.e., time t1. In addition, the nodding action just after the speech starts may not be a nodding action for that speech, and may be a nodding action for the speech just before that speech. Therefore, in order to properly divide the nodding action for speech a and the nodding action for the speech just before speech a, the evaluation value calculation unit 26 may also set the start timing of the evaluation period corresponding to speech a as the timing after a predetermined time has elapsed from the start timing of speech a, i.e., time t1.

[0069] In addition, the evaluation value calculation unit 26 sets the end timing (the second timing) of the evaluation period for speech a as the timing, i.e., time t3, after a predetermined time has elapsed from the end timing of speech a, i.e., time t2. Here, the predetermined time is preferably in the range of 5 seconds to 15 seconds, and is set to 15 seconds in this embodiment.

[0070] Then, the evaluation value calculation unit 26 calculates the evaluation value for speech a by summing up the f(p) values in the evaluation period corresponding to speech a.

[0071] (Statement b)

[0072] The evaluation value calculation unit 26 sets the start timing of the evaluation period corresponding to the statement b to time t4 by the same method.

[0073] On the other hand, according to Figure 8 , before the elapse of the above-mentioned predetermined time from time t5 which is the end timing of the statement b, the statement c starts. Therefore, when setting the end timing of the evaluation period corresponding to the statement b to the timing after the elapse of the above-mentioned predetermined time from time t5 in the same way as the end timing of the evaluation period corresponding to the statement a, there is a possibility of treating the nodding motion for the statement c as the nodding motion for the statement b. Therefore, in this case, the evaluation value calculation unit 26 sets the end timing of the evaluation period corresponding to the statement b to time t6 when the statement c starts.

[0074] In Figure 8 's example, the f(p) value during the process of the statement b is extremely low, but a relatively large f(p) value is observed just at the start of the statement c. The relatively large f(p) value approximately considered immediately after time t6 may be due to the statement c rather than the statement b. Therefore, by setting the end timing of the evaluation period corresponding to the statement b to time t6 when the statement c starts as described above, over-evaluation of the statement b is avoided.

[0075] Then, the evaluation value calculation unit 26 calculates the evaluation value for the statement b by summing up the f(p) values in the evaluation period corresponding to the statement b.

[0076] (Statement c)

[0077] The evaluation value calculation unit 26 sets the start timing and the end timing of the evaluation period corresponding to the statement c by the same method as the statement b, sums up the f(p) values in the evaluation period corresponding to the statement c, and thus calculates the evaluation value for the statement c.

[0078] (Statement d)

[0079] The evaluation value calculation unit 26 sets the start timing and the end timing of the evaluation period corresponding to the statement d by the same method as the statement a, sums up the f(p) values in the evaluation period corresponding to the statement d, and thus calculates the evaluation value for the statement d.

[0080] Then, the evaluation value calculation unit 26 associates the statement detected by the statement detection unit 22 with the start time of the statement, the voice data, and the evaluation value for the statement as shown in Figure 9 , and stores them as evaluation data in the data storage unit 21. The evaluation value for a statement is a powerful indicator representing the importance of the statement.

[0081] Then, the evaluation value output unit 27 outputs evaluation data by a desired method.

[0082] By referring to the output evaluation data, multiple participants 2 can simply obtain, in a short time, voice data of highly evaluated statements considered important in the communication. Therefore, for the participants 2 who are going to create a record of the communication, by preferentially listening to and viewing the voice data of highly evaluated statements, they can review the content of the communication in a shorter time and create an accurate record of the meeting in a shorter time.

[0083] Hereinafter, with reference to Figure 10 , the operation of the speech evaluation system 1 will be described.

[0084] S100:

[0085] First, the evaluation device 4 determines whether a communication composed of multiple participants 2 has started. If it is determined that the communication has not started (S100: No), the evaluation device 4 repeats S100. On the other hand, if it is determined that the communication has started (S100: Yes), the evaluation device 4 proceeds to S110. For example, when the communication between the evaluation device 4 and multiple wearable terminals 3 is established, the evaluation device 4 can determine that the communication has started.

[0086] S110:

[0087] Next, the data transceiver 20 receives the transmission data 14a from the multiple wearable terminals 3 and stores it in the data storage unit 21.

[0088] S120:

[0089] Next, the evaluation device 4 determines whether the communication composed of multiple participants 2 has ended. If it is determined that the communication has not ended (S120: No), the evaluation device 4 returns the process to S110. On the other hand, if it is determined that the communication has ended (S120: Yes), the evaluation device 4 proceeds to S130. For example, when the communication between all the wearable terminals 3 in communication with the evaluation device 4 and the evaluation device 4 is disconnected, the evaluation device 4 can determine that the communication has ended.

[0090] S130:

[0091] Next, the speech detection unit 22 refers to the transmission data 14a stored in the data storage unit 21, detects the speech in the communication, and determines the wearable terminal 3 corresponding to the detected speech.

[0092] S140:

[0093] Next, during the speech, the speech period detection unit 23 detects the start timing and end timing of each speech detected by the speech detection unit 22.

[0094] S150:

[0095] Next, the approval ratio calculation unit 24 calculates the approval ratio at every predetermined time interval.

[0096] S160:

[0097] Next, the emphasis processing unit 25 performs emphasis processing on the approval ratio calculated by the approval ratio calculation unit 24 to emphasize the high or low of the approval ratio.

[0098] S170:

[0099] Next, the evaluation value calculation unit 26 sets an evaluation period corresponding to each speech detected by the speech detection unit 22, and calculates the evaluation value for each speech.

[0100] S180:

[0101] Then, the evaluation value output unit 27 outputs evaluation data by a desired method.

[0102] The preferred embodiments of the present invention have been described above, but the above embodiments have the following features.

[0103] That is, in the communication composed of multiple participants 2, the speech evaluation system 1 that obtains the evaluation value for each speech includes a plurality of wearable terminals 3, a speech detection unit 22, a speech period detection unit 23, and an evaluation value calculation unit 26.

[0104] A plurality of wearable terminals 3 are provided with sensors 16, which are worn on each of a plurality of participants 2, and each sensor 16 includes at least a microphone 11 (sound collection unit). A speech detection unit 22 detects a speech during communication based on the output values of the microphones 11 of the plurality of wearable terminals 3, and determines the wearable terminal 3 corresponding to the detected speech. A speech period detection unit 23 detects the start timing and the end timing of each speech detected by the speech detection unit 22. An evaluation value calculation unit 26 calculates, for each speech detected by the speech detection unit 22, an evaluation value for the speech based on the output values of the acceleration sensors 12 of the wearable terminals 3 other than the wearable terminal 3 corresponding to the speech during an evaluation period (speech evaluation target period) from a first timing after the start timing of the speech and before the end timing to a second timing after the end timing of the speech. According to the above structure, in addition to the reactions of the audience during the speech, the reactions of the audience that occur after the speech with a delay are also reflected in the calculation of the evaluation value for the speech, so that the evaluation value for each speech can be accurately obtained.

[0105] In addition, the second timing is set to a timing at which a predetermined time has elapsed since the end timing of the corresponding speech. For example, refer to Figure 8 times t3 and t10. According to the above structure, the operation required to set the second timing is simplified, so that the second timing can be set at low cost.

[0106] In addition, the second timing is set to the start timing of another speech following the corresponding speech. For example, refer to Figure 8 times t6 and t8. According to the above structure, the evaluation value can be calculated excluding the reactions to other speeches, so that the evaluation value for the corresponding speech can be accurately obtained.

[0107] In addition, the second timing is set to a timing at which a predetermined time has elapsed since the end timing of the corresponding speech (refer to times t3 and t10). When other speeches (speech c, speech d) start after the corresponding speech before a predetermined time has elapsed since the end timing of the corresponding speech, the second timing is set to the start timing of the other speech following the corresponding speech. For example, refer to Figure 8 times t6 and t8. According to the above structure, when other speeches do not start after the corresponding speech before a predetermined time has elapsed since the end timing of the corresponding speech, the second timing can be set at low cost, and when other speeches start after the corresponding speech before a predetermined time has elapsed since the end timing of the corresponding speech, the evaluation value can be calculated excluding the reactions to other speeches, so that the evaluation value for the corresponding speech can be accurately obtained.

[0108] When the output value of the acceleration sensor 12 indicates that the participant wearing the corresponding wearable terminal 3 makes a head longitudinal swinging motion, the evaluation value calculation unit 26 calculates the evaluation value in a way that increases the evaluation value for the corresponding speech. That is, since the head longitudinal swinging motion can be regarded as an approval behavior, the corresponding speech can be regarded as a relatively high evaluation.

[0109] The above-described embodiment can be changed in the following manner.

[0110] In the above-described embodiment, the approval ratio calculation unit 24 extracts the vertical component value of the acceleration data and detects the nodding motion of the participant 2 based on the extracted vertical component value. However, instead of this operation, or in addition to this, the horizontal component value of the acceleration data can be extracted, and based on the extracted horizontal component value, the motion of the participant 2 making a head lateral swinging motion, that is, a rejecting motion, can be detected. The head lateral swinging motion is a motion in contrast to the nodding motion, that is, the head longitudinal swinging motion, and implies a negative or disagreeing meaning display for the speech. In this case, the approval ratio calculation unit 24 can also calculate the approval ratio in a way that cancels out the nodding motion and the rejecting motion. Thus, for example, when the number of participants 2 participating in the communication is 10, and in a certain time interval, 8 of them make nodding motions and the remaining two make rejecting motions, the approval ratio calculation unit 24 can calculate the approval ratio in this certain time interval as (8 - 2) / 10 = 0.6. In short, when the output value of the acceleration sensor 12 indicates that the participant wearing the corresponding wearable terminal 3 makes a head lateral swinging motion, the evaluation value calculation unit 26 can calculate the evaluation value in a way that decreases the evaluation value for the corresponding speech.

[0111] In the above-described embodiment, each wearable terminal 3 is equipped with an acceleration sensor 12, and the approval ratio calculation unit 24 calculates the approval ratio based on the output value of the acceleration sensor 12 of each wearable terminal 3. However, the acceleration sensor 12 can be omitted. In this case, the approval ratio calculation unit 24 calculates the approval ratio based on the output value of the microphone 11 of each wearable terminal 3. For example, when the microphone 11 of each wearable terminal 3 picks up a sound indicating approval such as "indeed", "certainly", "exactly so", etc., the approval ratio calculation unit 24 can regard this sound as an approval expression equivalent to the nodding motion and calculate the approval ratio.

[0112] In addition, the evaluation device 4 can also be constructed on a cloud system, and each wearable terminal 3 communicates with the evaluation device 4 via the Internet. In addition, the information processing performed by the evaluation device 4 can also be distributed and processed by multiple devices.

[0113] In addition, for example, as Figure 7As shown above, in the above-described embodiment, when the emphasis processing unit 25 performs emphasis processing on the approval ratio calculated by the approval ratio calculation unit 24 to emphasize the level of the approval ratio, a monotonically increasing function is used. However, instead of this function, as Figure 11 shown, when the emphasis processing unit 25 performs emphasis processing on the approval ratio calculated by the approval ratio calculation unit 24 to emphasize the level of the approval ratio, a step function represented by the following formula (2) may also be used.

[0114]

Formula 2

[0115]

[0116] In the above example, the program is stored using various types of non-transitory computer readable media and can be provided to a computer. Non-transitory computer readable media include various types of tangible storage media. Examples of non-transitory computer readable media include magnetic recording media (such as flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (such as magneto-optical disks). Examples of non-transitory computer readable media also include CD-ROM (Read Only Memory), CD-R, CD-R / W, semiconductor memories (such as, including mask ROM). Examples of non-transitory computer readable media also include PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM (random access memory)). In addition, the program can also be provided to a computer through various types of transitory computer readable media. Examples of transitory computer readable media include electrical signals, optical signals, and electromagnetic waves. Transitory computer readable media can provide the program to a computer via wired communication lines such as wires and optical fibers or wireless communication lines.

[0117] Either the evaluation device 4 can perform a part of the functions possessed by each wearable terminal 3, or any wearable terminal 3 can perform a part of the functions possessed by the evaluation device 4.

[0118] Based on the disclosure described as such, it is obvious that the embodiments of the present disclosure can be changed in various ways. Such changes should not be regarded as departing from the spirit and scope of the present disclosure, and all such modifications that are obvious to those skilled in the art are intended to be included within the scope of the appended claims.

Claims

1. A speech evaluation system that calculates an evaluation value for each speech in a communication composed of multiple participants, wherein, The speech evaluation system includes: A plurality of wearable terminals, each worn by one of the plurality of participants, each having at least a sensor including a sound collection unit and an acceleration sensor; A speech detection unit that detects a speech in the communication based on the output value of the sound collection unit of the plurality of wearable terminals, and determines the wearable terminal corresponding to the detected speech; A speech period detection unit that, for each speech detected by the speech detection unit, detects the start timing and the end timing of the speech; An approval ratio calculation unit that determines whether there is a nodding action based on the acceleration data from the acceleration sensor, and then calculates the approval ratio for a predetermined time interval, and An evaluation value calculation unit that, for each speech detected by the speech detection unit, calculates an evaluation value for the speech based on the output value of the sensor of the wearable terminals other than the wearable terminal corresponding to the speech during the speech evaluation target period and the approval ratio calculated by the approval ratio calculation unit. The speech evaluation target period is a period from a first timing after the start timing of the speech and earlier than the end timing to a second timing after the end timing of the speech; When the acceleration sensor detects an action of the participant wearing the corresponding wearable terminal swinging the head longitudinally, the evaluation value calculation unit calculates the evaluation value in a way that increases the evaluation value for the corresponding speech; The speech detection unit detects the presence of a speech in the communication based on sound data exceeding a predetermined value.

2. The speech evaluation system according to claim 1, wherein, The second timing is set to a timing after a predetermined time has elapsed from the end timing of the corresponding speech.

3. The speech evaluation system according to claim 1, wherein, The second timing is set to the timing when another speech starts after the corresponding speech.

4. The speech evaluation system according to claim 1, wherein, The second timing is set to a timing after a predetermined time has elapsed from the end timing of the corresponding speech. When another speech starts after the corresponding speech before a predetermined time has elapsed from the end timing of the corresponding speech, the second timing is set to the timing when another speech starts after the corresponding speech.

5. The speech evaluation system according to claim 1, wherein, When the acceleration sensor detects an action of the participant wearing the corresponding wearable terminal swinging the head laterally, the evaluation value calculation unit calculates the evaluation value in a way that decreases the evaluation value for the corresponding speech.

6. A speech evaluation method that calculates an evaluation value for each speech in a communication composed of multiple participants, A plurality of wearable terminals are worn by each of the multiple participants, and each of the plurality of wearable terminals has at least a sensor including a sound collection unit and an acceleration sensor, The speech evaluation method includes: Detect a speech in the communication based on the output value of the sound collection unit of the plurality of wearable terminals, and determine the wearable terminal corresponding to the detected speech; For each detected speech, detect the start timing and the end timing of the speech; Determine whether there is a nodding action based on the acceleration data from the acceleration sensor, and then calculate the approval ratio for a predetermined time interval, and For each detected speech, a evaluation value for the speech is calculated based on the output value of the sensor of the wearable terminal other than the wearable terminal corresponding to the speech and the calculated approval ratio during the speech evaluation object period, where the speech evaluation object period is a period from the first timing after the start timing of the speech and earlier than the end timing to the second timing after the end timing of the speech. In the case where the acceleration sensor detects an action of a participant wearing the corresponding wearable terminal swinging the head longitudinally, the evaluation value is calculated in a manner that increases the evaluation value for the corresponding speech. During the process of detecting the speech, the presence of a speech in the communication is detected based on voice data exceeding a predetermined value.

7. A computer recording medium, wherein, The computer recording medium records a program that causes a computer to execute the speech evaluation method according to claim 6.

Citation Information

Patent Citations

  • Conversation analysis device, conversation analysis system, conversation analysis method and conversation analysis program

    JP2016103081A

  • Voice recognition device and method for vehicle

    US10621985B2

  • Speech recognition

    US20040243416A1

  • Conference system, conference server, and program

    US20190394247A1