Learning device, determination device, determination system, learning method, determination method, trained model, program and learning model generation method
A learning device and system analyze speech features to determine the entertainingness of speaking styles in anecdotal talk, addressing the lack of quantitative evaluation in existing technologies.
Patent Information
- Application Number
- JP2024021226
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-15
- Publication Date
- 2025-08-27
AI Technical Summary
Existing technologies lack the ability to quantitatively evaluate and determine the entertainingness of a speaking style in anecdotal talk, particularly in oral performances like stand-up comedy, beyond vague and qualitative descriptions.
A learning device and system that utilize feature information such as duration and number of pronunciations, pauses, and speech rates in punchline and introductory sections to learn a judgment model that evaluates the entertainingness of a speaking style.
Enables the determination of whether a speaking style is interesting by analyzing specific features of speech patterns, providing a quantitative assessment of entertainingness.
Smart Images

Figure 2025125270000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a learning device, a determination device, a determination system, a learning method, a determination method, a trained model, a program, and a training model generation method. [Background technology]
[0002] Conventionally, there has been a technique for quantitatively evaluating the entertainment value of oral performances such as stand-up comedy and comedy sketches (see, for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-22118 Summary of the Invention [Problem to be solved by the invention]
[0004] Anecdotal talk is when a person tells a humorous story about an event. A good anecdote talk can make the listener laugh with a punch line at the end. However, techniques for making anecdotes more entertaining have been limited to vague and quantitative content, such as "putting emotion into it" or "telling specific events." As a result, it has been difficult to actually incorporate these techniques and achieve an entertaining style of speaking. The technology in Patent Document 1 evaluates the entertainment value of a performance, but does not determine whether the way the anecdote talk is spoken is entertaining.
[0005] The present invention has been made in view of the above circumstances, and provides a technique that makes it possible to determine whether a speaking style is interesting. [Means for solving the problem]
[0006] One aspect of the present invention is a learning device that includes a learning unit that uses training data including feature information that represents characteristics of a speaker's speaking style, obtained based on the duration and number of pronunciations of each of a plurality of speech sections that belong to a punch line part of the speaker's speech, the duration of each of a plurality of pauses that belong to the punch line part, the duration and number of pronunciations of each of a plurality of speech sections that belong to an introductory part of the speech, and the duration of each of a plurality of pauses that belong to the introductory part, and an evaluation score that represents the entertainingness of the speaker's speaking style, as input of the feature information obtained from the speech of a person to be evaluated, and learns a judgment model that outputs an evaluation score for the entertainingness of the speaker's speaking style.
[0007] One aspect of the present invention is the above-mentioned learning device, wherein the feature information includes information on at least the first feature among a first feature including the average duration of multiple pauses belonging to the punchline section, the average number of pronunciations in multiple speech sections belonging to the punchline section, the average duration of multiple pauses belonging to the introduction section, and the average number of pronunciations in multiple speech sections belonging to the introduction section; a second feature including the speech rate of the punchline section and the speech rate of the introduction section; and a third feature including the average duration of multiple speech sections belonging to the punchline section and the average duration of multiple speech sections belonging to the introduction section.
[0008] One aspect of the present invention is a judgment device that includes a judgment unit that inputs feature information that represents the characteristics of a speaker's speaking style, which is obtained based on the duration and number of pronunciations of each of multiple speech sections that belong to a punch line part of the speaker's speech, the duration of each of multiple pauses that belong to the punch line part, the duration and number of pronunciations of each of multiple speech sections that belong to an introductory part of the speech, and the duration of each of multiple pauses that belong to the introductory part, and uses a judgment model that outputs an evaluation score that represents the entertainingness of the speaker's speaking style, and obtains an evaluation score for the entertainingness of the speaker's speaking style that corresponds to the feature information obtained from the speech of the person to be evaluated.
[0009] One aspect of the present invention is the above-mentioned determination device, wherein the feature information includes information on at least the first feature among a first feature including an average duration of multiple pauses belonging to the punchline section, an average number of pronunciations in multiple speech sections belonging to the punchline section, an average duration of multiple pauses belonging to the introduction section, and an average number of pronunciations in multiple speech sections belonging to the introduction section; a second feature including a speech rate of the punchline section and an utterance rate of the introduction section; and a third feature including an average duration of multiple speech sections belonging to the punchline section and an average duration of multiple speech sections belonging to the introduction section.
[0010] One aspect of the present invention is the above-mentioned determination device, further comprising a preprocessing control unit that detects multiple pauses and multiple speech sections based on the volume of the voice indicated by the audio data of the speech of the person to be evaluated, defines a predetermined number of the last speech sections and pauses among the detected multiple speech sections as punch lines, and defines the section from the start of the speech to the punch line as an introduction line, and acquires the feature information based on the pauses and speech sections included in the detected punch line and introduction line, respectively, and the number of pronunciations in each of the speech sections obtained based on the speech recognition result of the audio data.
[0011] One aspect of the present invention is a judgment system comprising: a learning unit that uses training data including feature information representing characteristics of a speaker's speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of the speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and an evaluation score representing the entertainingness of the speaker's speaking style; and a judgment unit that uses the trained judgment model to obtain an evaluation score corresponding to the feature information obtained from the speaker's speaking style.
[0012] One aspect of the present invention is a learning method having a learning step of using training data including feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line part of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line part, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory part of the speech, and the duration of each of a plurality of pauses belonging to the introductory part, and an evaluation score representing the entertainingness of the speaking style of the speaker, as input of the feature information obtained from the speech of a person to be evaluated, and training a judgment model that outputs an evaluation score for the entertainingness of the speaking style of the person to be evaluated.
[0013] One aspect of the present invention is a judgment method having a judgment step of inputting feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and using a judgment model that outputs an evaluation score representing the entertainingness of the speaking style of the speaker, and obtaining an evaluation score for the entertainingness of the speaking style of the person to be evaluated that corresponds to the feature information obtained from the speech of the person to be evaluated.
[0014] One aspect of the present invention is a trained model trained using training data including feature information representing characteristics of a speaker's speaking style, obtained based on the duration and number of pronunciations of each of multiple speech sections belonging to the punch line part of the speaker's speech, the duration of each of multiple pauses belonging to the punch line part, the duration and number of pronunciations of each of multiple speech sections belonging to the introductory part of the speech, and the duration of each of multiple pauses belonging to the introductory part, and an evaluation score representing the entertainingness of the speaker's speaking style, and the trained model is used to input the feature information obtained from the speech of a person to be evaluated into a computer and execute a process of outputting an evaluation score for the entertaining speaking style of the person to be evaluated.
[0015] One aspect of the present invention is a program for causing a computer to function as the learning device described above.
[0016] One aspect of the present invention is a program for causing a computer to function as the above-described determination device.
[0017] One aspect of the present invention is a learning model generation method that includes a learning step of using training data including feature information of a speaker's speaking style, which is the average length of multiple pauses in the speaker's speech, and an evaluation score representing the interestingness of the speaker's speaking style, to input the feature information obtained from the speech of a person to be evaluated, and to learn a judgment model that outputs an evaluation score for the interestingness of the speaker's speaking style.
[0018] One aspect of the present invention is a learning model generation method including a learning step of using training data including feature information of a speaker's speaking style, which is the average of multiple pause lengths belonging to the introductory part of the speaker's speaking style, and an evaluation score representing the interestingness of the speaker's speaking style, to input the feature information obtained from the speech of a person to be evaluated, and to learn a judgment model that outputs an evaluation score for the interestingness of the speaker's speaking style.
[0019] One aspect of the present invention is the above-mentioned learning model generation method, wherein the feature information further includes a variance of the speaking rate of a plurality of speech sections in the speaker's speech.
[0020] One aspect of the present invention is the above-mentioned learning model generation method, wherein the feature information further includes a variance of the speaking rate of a plurality of speech sections belonging to an introductory part of the speaker's speech. [Effects of the Invention]
[0021] The present invention makes it possible to determine whether a person's speaking style is interesting. [Brief explanation of the drawings]
[0022] [Figure 1] FIG. 10 is a diagram showing information obtained from audio data of an episode talk in one embodiment of the present invention. [Figure 2] FIG. 10 is a diagram showing analytical information obtained from audio data of episode talk in the same embodiment. [Figure 3]FIG. 10 is a diagram showing the analysis results of speaking style according to the embodiment. [Figure 4] FIG. 2 is a schematic block diagram showing the system configuration of a determination system according to the embodiment. [Figure 5] FIG. 2 is a schematic block diagram illustrating an example of the functional configuration of a terminal device according to the embodiment. [Figure 6] FIG. 2 is a schematic block diagram illustrating an example of the functional configuration of a learning device according to the embodiment. [Figure 7] 10A and 10B are diagrams illustrating examples of teacher data and preprocessed teacher data according to the embodiment. [Figure 8] FIG. 10 is a diagram showing an example of a radar chart display according to the embodiment. [Figure 9] FIG. 2 is a diagram illustrating an example of a determination model according to the embodiment. [Figure 10] FIG. 2 is a schematic block diagram illustrating an example of a functional configuration of a determination device according to the embodiment. [Figure 11] 10 is a flowchart illustrating an example of processing performed by the learning device according to the embodiment. [Figure 12] 10 is a flowchart showing an example of processing performed by the determination device according to the embodiment. [Figure 13] FIG. 2 is a diagram illustrating an outline of a hardware configuration example of an information processing device applied to the embodiment. [Figure 14] FIG. 10 is a schematic block diagram showing a modified example of the determination device according to the embodiment. [Figure 15] FIG. 10 is a schematic block diagram showing a modified example of the determination device according to the embodiment. [Figure 16] 10A and 10B are diagrams showing experimental results using the determination system according to the embodiment. [Figure 17] 10A and 10B are diagrams showing experimental results using the determination system according to the embodiment. [Figure 18] 10A and 10B are diagrams showing experimental results using the determination system according to the embodiment. [Figure 19] 10A and 10B are diagrams showing experimental results using the determination system according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0023] An embodiment of the present invention will be described below with reference to the drawings. A determination system according to an embodiment of the present invention determines whether a speaking style in episode talk is interesting. First, the features used by the determination system of this embodiment to determine whether a speaking style in episode talk is interesting will be described. Hereinafter, a speaking style in episode talk will be simply referred to as "speaking style," an "interesting speaking style in episode talk" will be referred to as "interesting speaking style," and an "uninteresting speaking style in episode talk" will be referred to as "uninteresting speaking style."
[0024] FIG. 1 is a diagram showing information that the determination system of this embodiment acquires from audio data of episode talk to analyze speaking styles. Audio data can also be acquired from video (image) data of episode talk, as shown in FIG. 1. In episode talk, time periods in which the speaker is speaking and time periods in which the speaker is not speaking are repeated alternately. Each time period in which the speaker is speaking is referred to as a "speech period," and each time period in which the speaker is not speaking (non-speech period) is referred to as a "pause."
[0025] The determination system of this embodiment performs speech recognition on the audio data of the episode talk and transcribes it to obtain text data of the speaker's speech (speech content) during the episode talk. Furthermore, the determination system detects speech sections and pauses within the episode talk based on the speaker's voice volume indicated by the audio data. The determination system calculates the duration of each speech section and each pause based on the playback positions where the detected speech sections and pauses switch. The playback positions are expressed as the elapsed time from an arbitrary position, such as the beginning of the audio data, as a reference position.
[0026] FIG. 2 is a diagram showing analysis information acquired by the determination system of this embodiment from audio data of episode talk. The determination system monitors the sound pressure of the audio data in chronological order from the start of the episode talk, and detects a pause when the sound pressure remains below a threshold continuously for a period longer than a predetermined time T1 (e.g., 0.2 seconds). The determination system then detects the end of a pause, i.e., the start of a speech period, when the sound pressure exceeds the threshold after the start of the pause. The determination system repeats this detection of pauses and speech periods. The determination system detects the end of episode talk when the sound pressure remains below the threshold for a period longer than a predetermined time T2 (T2>T1). This allows the determination system to obtain the speech periods and pauses from the start to the end of the episode talk in the audio data.
[0027] If the number of speech sections included in an episode talk is N, these N speech sections will be referred to as speech sections #1 to #N in the order of their appearance. Furthermore, when k is an integer between 1 and N-1, the section between speech section #k and speech section #(k+1) will be referred to as interval #k. If the start timing of speech section #n (n is an integer between 1 and N) is playback position t(2n-1) and the end timing is playback position t(2n), the duration of speech section #n is t(2n)-t(2n-1). Furthermore, the start timing of interval #k is playback position t(2k), the end timing is playback position t(2k+1), and the duration of interval #k is t(2k+1)-t(2k).
[0028] Furthermore, the determination system obtains character data #n corresponding to each speech section #n from the character data of the speech recognition result. The determination system obtains the number of pronunciations in the speech section #n from this character data #n. In this embodiment, the number of moras is used as the number of pronunciations, but the number of characters when the speech content is expressed in kana may also be used. The determination system obtains the speaking rate in the speech section #n by dividing the number of pronunciations obtained from the character data #n by the duration of the speech section #n.
[0029] On the other hand, analysis of multiple episode talks has revealed that the punch line that makes listeners laugh is often found in the last three (or two) speech sections. Therefore, the determination system determines the time section from speech section #1 to speech section #(N-3) as the introductory part of the episode talk, and the time section from speech section #(N-3) to speech section #N thereafter as the punch line of the episode talk.
[0030] Figure 3 shows the results of an analysis of funny and unfunny speaking styles using actual episode talks. If the listeners laughed at the end of an episode talk, it was classified as funny, and if they did not laugh, it was classified as funny. Then, for 60 funny and 60 unfunny speaking styles, the average and variance of the duration between the entire episode talk, the punch line, and the introduction, the duration of each speech segment, the number of pronunciations in each speech segment, and the speaking rate were calculated. Figure 3 shows the average of the analysis results.
[0031] The analysis results revealed the following tendencies in interesting speaking styles: (1) The overall pauses (duration) were long (approximately 0.87 seconds), and the overall variation in pause duration was large (variance 0.23). (2) Overall, the number of pronunciations in each utterance section was small (average 10.73 moras) and the variation was small (variance 37). (3) Overall, the speech rate is slow (approximately 8.7 mora / second). (4) The time between punch lines is longer than the time between introduction lines (+0.22 seconds). (5) The speech rate of the punch line is faster than that of the introduction line (-0.44 pronunciations per second), and the number of pronunciations in the punch line is greater than that of the introduction line (+1.49 pronunciations per second).
[0032] On the other hand, in uninteresting speech, the pauses are shorter overall. Also, in uninteresting speech, the duration of each utterance segment is longer in the introduction than in the punch line, which is the opposite of interesting speech. Similarly, in uninteresting speech, the number of pronunciations in each utterance segment is greater in the introduction than in the punch line, which is the opposite of interesting speech.
[0033] Therefore, the determination system of this embodiment uses the average duration (seconds) between punchline sections and between introductory sections, and the average number of pronunciations (mora) between the speech sections of the punchline section and the speech section of the introductory section as feature quantities for determining whether a speech style is unfunny or funny. To further improve the determination accuracy, the average speech rate (mora / second) between the punchline section and the introductory section, and the variance of the speech rate for the introductory section are also used as feature quantities. Furthermore, because the speech rate can be calculated based on the number of pronunciations and the duration, the average duration (seconds) between the speech sections of the punchline section and the introductory section may be used instead of the average speech rate between the punchline section and the introductory section.
[0034] Alternatively, the judgment system of this embodiment may use as features the average and variance of the length of pauses in the entire episode talk including the punchline and introduction, the average and variance of the number of utterances in the speech section in the entire episode talk, the average and variance of the speaking rate in the entire episode talk, and the average and variance of the length of the speech section in the entire episode talk.
[0035] Next, the configuration of the determination system of this embodiment will be described. Fig. 4 is a schematic block diagram showing the system configuration of a determination system 100 according to one embodiment of the present invention. The determination system 100 is used to determine whether a speaking style is interesting or not based on feature information (hereinafter referred to as "speech feature information") that is obtained from data of utterances in episode talk and that represents the characteristics of the speaking style. More specifically, the utterance feature information may be used to obtain an evaluation score that quantitatively represents the interestingness of the speaking style.
[0036] The determination system 100 includes a terminal device 10, a learning device 20, and a determination device 30. Although only one terminal device 10 is shown in FIG. 4, the number of terminal devices 10 is arbitrary. The terminal device 10 and the determination device 30 are communicably connected via a network 70. The learning device 20 and the determination device 30 may also be communicably connected via the network 70. The network 70 may be a network using wireless communication or a network using wired communication. The network 70 may be configured using, for example, the Internet or a local area network (LAN). The network 70 may also be configured by combining multiple networks.
[0037] FIG. 5 is a schematic block diagram showing a specific example of the functional configuration of the terminal device 10. In FIG. 5, only functional blocks related to this embodiment are extracted and shown. The terminal device 10 is configured using information devices such as a smartphone, tablet, personal computer, or dedicated device. The terminal device 10 includes a communication unit 11, an input unit 12, a display unit 13, an audio output unit 14, a storage unit 15, and a control unit 16.
[0038] The communication unit 11 is a communication device. The communication unit 11 may be configured as, for example, a network interface. The communication unit 11 communicates data with other devices via the network 70 in accordance with the control of the control unit 16. The communication unit 11 may be a device that performs wireless communication or a device that performs wired communication.
[0039] The input unit 12 is a keyboard, a mouse, a button, a touch panel, or the like, and receives information input by user operation.
[0040] The display unit 13 outputs information in a form that can be recognized by the user. The display unit 13 may be, for example, an image display device such as a liquid crystal display or an organic EL (Electro Luminescence) display. The display unit 13 may also be an interface for connecting an image display device to the terminal device 10. In this case, the display unit 13 generates a video signal for displaying image data and outputs the video signal to the image display device connected to the display unit 13. The display unit 13 may also be configured as a touch panel integrated with the input unit 12.
[0041] The audio output unit 14 is a device that outputs sound, such as a speaker. The audio output unit 14 may be an interface for connecting an audio output device, such as a speaker or headphones, to the terminal device 10. In this case, the audio output unit 14 generates an audio signal for reproducing audio data and outputs the audio signal to the audio output device connected to the audio output unit 14.
[0042] The storage unit 15 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 15 stores data used by the control unit 16. The storage unit 15 stores data required when the control unit 16 performs processing.
[0043] The control unit 16 is configured using a processor such as a CPU (Central Processing Unit) and a memory (main storage device). The control unit 16 functions when the processor executes a program. All or part of the functions of the control unit 16 may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs: Solid State Drives), as well as storage devices such as hard disks and semiconductor storage devices built into computer systems. The program may be transmitted via a telecommunications line.
[0044] The control unit 16 controls the terminal device 10 in response to user operations and information received from the determination device 30. For example, the control unit 16 may output, from the audio output unit 14, audio data of episode talk instructed to be played by the user operating the input unit 12. If the playback instruction targets video data of episode talk, the control unit 16 may display the instructed video data on the display unit 13 and output audio data included in the video data from the audio output unit 14. The control unit 16 also transmits information input by the user operating the input unit 12 to the determination device 30 using the communication unit 11. For example, such information includes audio data and video data of episode talk of a person being evaluated for their interesting speaking style. For example, when the communication unit 11 receives information transmitted from the determination device 30 via the network 70, the control unit 16 generates screen data based on the received information and displays the screen data on the display unit 13. Such screen data includes images and text representing the information transmitted from the determination device 30. For example, such information includes the evaluation results of the interestingness of the speaking style of the person to be evaluated, and a radar chart of speech feature information obtained from the audio data of the episode talk.
[0045] FIG. 6 is a schematic block diagram showing a specific example of the functional configuration of learning device 20. In FIG. 6, only functional blocks related to this embodiment are shown. Learning device 20 is configured using an information processing device such as a personal computer or a server device. Learning device 20 includes a communication unit 21, a memory unit 22, a control unit 23, an input unit 24, a display unit 25, and an audio output unit 26.
[0046] The communication unit 21 is a communication device. The communication unit 21 may be configured as, for example, a network interface. The communication unit 21 communicates data with other devices via the network 70 in accordance with the control of the control unit 23. The communication unit 21 may be a device that performs wireless communication or a device that performs wired communication.
[0047] The storage unit 22 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 22 stores data used by the control unit 23. The storage unit 22 may function as, for example, a teacher data storage unit 221, a preprocessed teacher data storage unit 222, and a trained model storage unit 223.
[0048] The training data storage unit 221 stores training data used in the learning process executed by the learning device 20. The training data stored in the training data storage unit 221 includes audio data or video data of episode talk and label information indicating an evaluation score that quantitatively represents the entertainingness of the speaking style of the episode talk. The evaluation score may be expressed as a binary value, such as 1 for entertaining speaking style and 0 for uninteresting speaking style, or may represent a score out of 100 or a level out of 5. However, for training data used to train a single determination model, evaluation scores are assigned using the same evaluation method and evaluation criteria. For example, the evaluation score may be 1 if the listener laughed at the end and 0 if they did not laugh. Furthermore, the evaluation score may be a score assigned to the episode talk by a judge or audience member at a program or event, or a value corresponding to the volume of laughter.
[0049] The preprocessed teacher data storage unit 222 stores preprocessed teacher data. The preprocessed teacher data is teacher data that includes information obtained by preprocessing teacher data. For example, the preprocessed teacher data may be data generated by adding speech feature information to teacher data, or data in which audio data or video data included in the teacher data is replaced with speech feature information. The speech feature information includes values of one or more types of feature parameters. In this embodiment, the speech feature information includes the following six feature parameters P1 to P6.
[0050] The feature parameter P1 is the average duration between punch lines, the feature parameter P2 is the average number of pronunciations in the speech sections of the punch line, and the feature parameter P3 is the speech rate in the speech sections of the punch line. The feature parameter P4 is the average duration between introduction lines, the feature parameter P5 is the average number of pronunciations in the speech sections of the introduction line, and the feature parameter P6 is the speech rate in the speech sections of the introduction line. All of the feature parameters P1 to P6 are used as explanatory variables. Note that the feature parameters P3 and P6 do not necessarily have to be included. Alternatively, the speech feature information may include a feature parameter P7 that is the average duration of the speech sections of the punch line and a feature parameter P8 that is the average duration of the speech sections of the introduction line, thereby making eight feature parameters the speech feature information. Alternatively, the feature parameter P7 may be used instead of the feature parameter P3, and the feature parameter P8 may be used instead of the feature parameter P6.
[0051] Alternatively, the speech feature information may include the following eight feature parameters P11 to P18. Feature parameter P11 is the average duration of the entire speech including the introduction and punchline, and feature parameter P12 is the variance of the duration of the entire speech. Feature parameter P13 is the average number of pronunciations in the entire speech section, and feature parameter P14 is the variance of the number of pronunciations in the entire speech section. Feature parameter P15 is the average overall speech rate, and feature parameter P16 is the variance of the overall speech rate. Feature parameter P17 is the average duration of the entire speech section, and feature parameter P18 is the variance of the duration of the entire speech section.
[0052] The trained model storage unit 223 stores a judgment model obtained by a learning process using the preprocessed teacher data stored in the preprocessed teacher data storage unit 222. The judgment model is a model that obtains an evaluation score for an interesting speaking style based on explanatory variables. The trained judgment model is also referred to as a trained model.
[0053] The control unit 23 is configured using a processor such as a CPU or GPU, and a memory. The control unit 23 functions as an information control unit 231, a preprocessing control unit 232, and a learning control unit 233 by the processor executing a program. All or part of the functions of the control unit 23 may be realized using hardware such as an ASIC, PLD, or FPGA. The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, and a semiconductor storage device (e.g., an SSD), and storage devices such as a hard disk or semiconductor storage device built into a computer system. The above program may be transmitted via a telecommunications line.
[0054] The information control unit 231 controls the input and output of information. For example, the information control unit 231 acquires training data from other devices (information processing devices or storage media) and records the training data in the training data storage unit 221. For example, the information control unit 231 transmits the trained model stored in the trained model storage unit 223 to another device (for example, the determination device 30).
[0055] The preprocessing control unit 232 generates preprocessed teacher data by performing a predetermined preprocessing on the teacher data. For example, the preprocessing control unit 232 may acquire speech feature information by performing a predetermined calculation on speech data included in the teacher data, and add the acquired speech feature information to the teacher data to generate preprocessed teacher data. Specifically, the preprocessing control unit 232 may perform the following process.
[0056] First, the preprocessing control unit 232 acquires time-series sound pressure information from the audio data of the episode talk. The preprocessing control unit 232 detects a pause when the sound pressure is continuously below a threshold for a period longer than the predetermined time T1. Furthermore, the preprocessing control unit 232 detects the end of the episode talk when the sound pressure is continuously below a threshold for a period longer than the predetermined time T1. The preprocessing control unit 232 defines the intervals between adjacent pauses and the interval between the final pause and the end of the episode talk as speech intervals. If the number of speech intervals in the punch line portion is a predetermined number M and the number of speech intervals detected from the audio data is N, the preprocessing control unit 232 defines the intervals from speech interval #1 to the (M+1)th speech interval from the end #(NM) as the introduction portion of the story, and the intervals from the subsequent pause #(NM) to the final speech interval #N as the punch line. M is, for example, 2 or 3, but is not limited to this. Here, a case where M=3 is used as an example.
[0057] The pre-processing control unit 232 obtains the duration of each pause and each speech section based on the playback positions of each pause and each start and end position of each speech section in the audio data. Furthermore, the pre-processing control unit 232 obtains character data #1 to #N of each speech section #1 to #N by performing voice recognition on the audio data, and calculates the number of pronunciations of each character data #1 to #N.
[0058] The pre-processing control unit 232 calculates the average duration of each of the punchline intervals #(N-3) to #(N-1) as the feature parameter P1. The pre-processing control unit 232 calculates the average number of pronunciations for each of the speech sections #(N-2) to #N in the punchline as the feature parameter P2. The pre-processing control unit 232 obtains the speech rate for each of the speech sections #(N-2) to #N, and calculates the average of the obtained speech rates as the feature parameter P3. The speech rate for speech section #n is calculated by dividing the number of pronunciations for speech section #n by the duration of speech section #n. Note that the pre-processing control unit 232 may also obtain the result of dividing the sum of the number of pronunciations for each of the speech sections #(N-2) to #N by the sum of the durations of each of the speech sections #(N-2) to #N, as the feature parameter P3. Furthermore, the preprocessing control unit 232 may determine the average duration of each of the speech sections #(N-2) to #N in the punch line part as feature parameter P7, which may be used in place of feature parameter P3 or may be added to feature parameters P1 to P3.
[0059] The pre-processing control unit 232 calculates the average duration of each of the speech sections #1 to #(N-4) in the introduction portion as feature parameter P4. The pre-processing control unit 232 calculates the average number of pronunciations of each of the speech sections #1 to #(N-3) in the introduction portion as feature parameter P5. The pre-processing control unit 232 calculates the speech rate of each of the speech sections #1 to #(N-3) and sets the average of the calculated speech rates as feature parameter P6. The pre-processing control unit 232 may also calculate the result of dividing the sum of the number of pronunciations of each of the speech sections #1 to #(N-3) by the sum of the durations of each of the speech sections #1 to #(N-3) as feature parameter P6. The pre-processing control unit 232 may also calculate the average duration of each of the speech sections #1 to #(N-3) in the introduction portion as feature parameter P8, which may be used in place of feature parameter P6 or may be added to feature parameter P4 to P6.
[0060] Alternatively, the pre-processing control unit 232 calculates the average duration of each of the overall pauses #1 to #(N-1) as feature parameter P11, and calculates the variance of the duration of each of the overall pauses #1 to #(N-1) as feature parameter P12. The pre-processing control unit 232 calculates the average number of pronunciations of each of the overall speech sections #1 to #N as feature parameter P13, and calculates the variance of the number of pronunciations of each of the overall speech sections #1 to #N as feature parameter P14. Furthermore, the pre-processing control unit 232 calculates the speech rate of each of the overall speech sections #1 to #N, and sets the average and variance of the calculated speech rates as feature parameter P15 and feature parameter P16, respectively. The pre-processing control unit 232 calculates the average duration of each of the overall speech sections #1 to #N as feature parameter P17, and calculates the variance of the duration of each of the overall speech sections #1 to #N as feature parameter P18. Furthermore, the preprocessing control unit 232 may use the variance of the speech rate of each of the speech sections #1 to #(N-3) in the introduction part instead of or in addition to the feature parameter P16.
[0061] In the following, an example will be described in which the speech feature information is feature parameters P1 to P6, but it may be feature parameters P1, P2, P4, and P5, feature parameters P1 to P8, feature parameters P1, P2, P4, P5, P7, and P8, or feature parameters P11 to P18. Furthermore, the speech feature information may be one or more feature parameters arbitrarily selected from the feature parameters P1 to P8 and P11 to P18.
[0062] In addition, the preprocessing control unit 232 generates radar chart display data for displaying radar charts of the feature parameters P1 to P6. The preprocessing control unit 232 adds the utterance feature information and the radar chart display data to the teacher data to generate preprocessed teacher data, and writes the preprocessed teacher data to the preprocessed teacher data storage unit 222.
[0063] The learning control unit 233 executes a learning process for a judgment model using the preprocessed teacher data stored in the preprocessed teacher data storage unit 222. Any type of artificial intelligence may be used for the judgment model. Any machine learning may be used for learning the judgment model. Specific examples of such learning processes include supervised learning for classification, such as neural networks, support vector machines, and random forests. The learning control unit 233 generates a judgment model for outputting an evaluation score for an interesting speaking style based on input utterance feature information, for example, by performing supervised learning. The learning control unit 233 records the generated judgment model as a learned model in the learned model storage unit 223. The learned model obtained by the learning control unit 233 may be transmitted to the judgment device 30 and recorded in a judgment model storage unit 321 of the judgment device 30, which will be described later.
[0064] The input unit 24 is a keyboard, a mouse, a button, a touch panel, or the like, and receives information input by user operation.
[0065] Display unit 25 outputs information in a form that can be recognized by the user. Display unit 25 may be an image display device such as a liquid crystal display or an organic EL display. Display unit 25 may also be an interface for connecting an image display device to learning device 20. In this case, display unit 25 generates a video signal for displaying image data and outputs the video signal to the image display device connected to it. Display unit 25 may also be configured as a touch panel integrated with input unit 24.
[0066] Audio output unit 26 is a device that outputs sound, such as a speaker. Audio output unit 26 may be an interface for connecting an audio output device, such as a speaker or headphones, to learning device 20. In this case, audio output unit 26 generates an audio signal for playing back audio data and outputs the audio signal to the audio output device connected to itself.
[0067] Figure 7 shows examples of teacher data and preprocessed teacher data. The teacher data shown in Figure 7(a) is data in which label information indicating an evaluation score is added to audio data or video data of an episode talk. The preprocessed teacher data shown in Figure 7(b) includes, in addition to the teacher data, values of each feature parameter P1 to P6 of the utterance feature information and radar chart display data that displays the values of each feature parameter P1 to P6 in a radar chart.
[0068] FIG. 8 is a diagram showing an example of a radar chart display. Each axis of the radar chart corresponds to a type of feature parameter P1 to P6 included in the speech feature information. The radar chart is represented by plotting the value of the feature parameter Pi (i is an integer between 1 and 6) on an axis corresponding to the type of the feature parameter Pi, and connecting the plot positions on adjacent axes with straight lines. The outermost scale values on each axis can be determined arbitrarily. The trend of an interesting speaking style and an uninteresting speaking style can be visualized by a figure represented by straight lines connecting the plot positions of the values of the feature parameters P1 to P6 on the radar chart.
[0069] Fig. 9 is a diagram showing an example of a judgment model. This figure shows a case where a DNN (Deep Neural Network) is used as the judgment model. The judgment model has an input layer that inputs the values of feature parameters P1 to P6 represented by a radar chart, and an output layer that outputs an evaluation score for an interesting speaking style.
[0070] FIG. 10 is a schematic block diagram showing a specific example of the functional configuration of the determination device 30. In FIG. 10, only functional blocks related to this embodiment are shown. The determination device 30 is configured using an information processing device such as a personal computer or a server device. The determination device 30 includes a communication unit 31, a storage unit 32, and a control unit 33.
[0071] The communication unit 31 is a communication device. The communication unit 31 may be configured as, for example, a network interface. The communication unit 31 communicates data with other devices via the network 70 in accordance with the control of the control unit 33. The communication unit 31 may be a device that performs wireless communication or a device that performs wired communication.
[0072] The storage unit 32 is configured using a storage device such as a magnetic hard disk drive or a semiconductor storage device. The storage unit 32 stores data used by the control unit 33. The storage unit 32 may function as a determination model storage unit 321, for example.
[0073] The judgment model storage unit 321 stores a judgment model used by the judgment unit 333 when performing the judgment process. The judgment model may be configured using information of a trained model generated in advance by a learning process, for example. Such a learning process may be executed by another device (for example, the learning device 20) or by the device itself (the judgment device 30). The judgment model does not necessarily have to be generated by a learning process. The judgment model may be configured using, for example, a lookup table that associates the value of each feature parameter with an evaluation value of an interesting speaking style, or may be configured in another manner.
[0074] The control unit 33 is configured using a processor such as a CPU and a memory. The control unit 33 functions as an information control unit 331, a preprocessing control unit 332, and a determination unit 333 by the processor executing a program. All or part of the functions of the control unit 33 may be realized using hardware such as an ASIC, a PLD, or an FPGA. The above program may be recorded on a computer-readable recording medium. Examples of computer-readable recording media include portable media such as flexible disks, magneto-optical disks, ROMs, CD-ROMs, and semiconductor storage devices (e.g., SSDs), as well as storage devices such as hard disks and semiconductor storage devices built into a computer system. The above program may be transmitted via a telecommunications line.
[0075] The information control unit 331 acquires audio data or video data of an episode talk being spoken by the person to be evaluated from another device such as the terminal device 10. The information control unit 331 also transmits information indicating the determination result obtained by the determination unit 333 to another device such as the terminal device 10. Such exchange of information between the information control unit 331 and another device may be performed by communication using the communication unit 31, for example.
[0076] The preprocessing control unit 332 acquires audio feature information by performing predetermined preprocessing on the audio data acquired by the information control unit 331 or the audio data extracted from video data. Furthermore, the preprocessing control unit 332 generates radar chart display data for the acquired audio feature information. The preprocessing performed by the preprocessing control unit 332 is similar to the preprocessing performed when generating a trained model used by the determination unit 333. In other words, the processing performed by the preprocessing control unit 332 on the audio data of episode talk is the same as the processing performed by the preprocessing control unit 232 on the audio data obtained from the training data. The preprocessing control unit 332 acquires audio feature information by performing such preprocessing, and generates radar chart display data for displaying the values of feature parameters P1 to P6 included in the audio feature information as a radar chart.
[0077] The determination unit 333 performs a determination process using the determination model stored in the determination model storage unit 321 and the speech feature information generated by the preprocessing control unit 332. The determination process determines whether the speaking style of the person to be evaluated is interesting. In this embodiment, the determination unit 333 obtains an evaluation score for the interesting speaking style.
[0078] Next, the operation of the determination system 100 will be described. FIG. 11 is a flowchart showing a specific example of processing by the learning device 20. First, the information control unit 231 acquires training data (step S101). The training data may be input by a user, acquired via communication from another information device such as the terminal device 10, or acquired from a recording medium connected to the learning device 20. Alternatively, the training data may be generated as follows: The control unit 23 acquires audio data of an episode talk instructed to be played by the user operating the input unit 24, and outputs the audio data from the audio output unit 26. Alternatively, the control unit 23 acquires video data of an episode talk instructed to be played by the user operating the input unit 24, displays the video data on the display unit 25, and outputs the audio data included in the video data from the audio output unit 26. When the user inputs an evaluation score for the interestingness of the speaking style via the input unit 12, the control unit 23 generates training data by associating the played audio data or video data with the input evaluation score.
[0079] The preprocessing control unit 232 performs predetermined preprocessing on each piece of training data to generate speech feature information consisting of the values of feature parameters P1 to P6 and radar chart display data representing the values of the feature parameters P1 to P6. The preprocessing control unit 232 adds the generated speech feature information and radar chart display data to the training data to generate preprocessed training data. The preprocessing control unit 232 writes the generated preprocessed training data to the preprocessed training data storage unit 222 (step S102).
[0080] The learning control unit 233 executes a learning process using the plurality of preprocessed teacher data stored in the preprocessed teacher data storage unit 222, and learns a judgment model that receives as input the values of the feature parameters P1 to P6 indicated by the speech feature information and outputs an evaluation value of the interestingness of the speaking style. The learning control unit 233 records the learned judgment model as a learned model in the learned model storage unit 223 (step S103).
[0081] In step S101, the information control unit 231 may receive preprocessed teacher data generated by an external device and write the preprocessed teacher data to the preprocessed teacher data storage unit 222. In this case, the preprocessing control unit 232 may not perform the process of step S102. Furthermore, the information control unit 231 may read a radar chart from the preprocessed teacher data and display it on the display unit 25 in response to an instruction input by the user to the input unit 24. For example, the information control unit 231 may read radar chart display data from preprocessed teacher data identified by information on audio data or video data input by the input unit 24, or from each preprocessed teacher data whose evaluation score falls within the numerical range input by the input unit 24, and display the radar chart display data on the display unit 25. Furthermore, the information control unit 231 may read feature parameters P1 to P6 from each preprocessed teacher data whose evaluation score falls within the numerical range input by the input unit 24, generate data displaying the average values of the read feature parameters P1 to P6 as a radar chart, and display the data on the display unit 25.
[0082] 12 is a flowchart showing a specific example of the processing of the determination device 30. First, the information control unit 331 acquires audio data or video data of an episode talk of the person to be evaluated from the terminal device 10 (step S201). Note that the information control unit 331 may read the audio data or video data of the person to be evaluated from a recording medium, or may read it from an information processing device such as a server computer connected via the network 70.
[0083] The pre-processing control unit 332 performs a predetermined pre-processing on the audio data acquired by the information control unit 331 or the audio data extracted from the acquired video data, and generates speech feature information consisting of the values of feature parameters P1 to P6 and display data of a radar chart showing the values of the feature parameters P1 to P6 (step S202). The judgment unit 333 performs judgment processing to obtain an output of an evaluation score for an interesting speaking style by inputting the values of the feature parameters P1 to P6 indicated by the speech feature information generated by the pre-processing control unit 332 into a judgment model (step S203).
[0084] The information control unit 331 transmits determination result information including information indicating the determination result of the determination process in step S203 and display data of the radar chart to the terminal device 10 (step S204). The control unit 16 of the terminal device 10 displays the evaluation score and radar chart display data indicated by the determination result information received from the determination device 30 on the display unit 13. The control unit 16 may output the evaluation score by voice from the voice output unit 14.
[0085] Furthermore, the information control unit 331 of the determination device 30 may acquire each preprocessed teacher data stored in the preprocessed teacher data storage unit 222 of the learning device 20 and write it to the storage unit 32. The information control unit 331 of the determination device 30 receives a numerical range of evaluation points input from the input unit 12 from the terminal device 10, and identifies preprocessed teacher data whose evaluation points fall within the received numerical range. The information control unit 331 may read audio data, video data, or radar chart display data from each identified preprocessed teacher data and transmit them to the terminal device 10. The control unit 16 of the terminal device 10 outputs the received audio data from the audio output unit 14. Alternatively, the control unit 16 displays the received video data on the display unit 13, and outputs the audio data included in the video data from the audio output unit 14. Furthermore, the control unit 16 displays the received radar chart display data on the display unit 13. Alternatively, the information control unit 331 of the determination device 30 may read out the feature parameters P1 to P6 from each of the identified preprocessed teacher data, generate display data for displaying the average values of the read out feature parameters P1 to P6 in a radar chart, and transmit the display data to the terminal device 10. The control unit 16 of the terminal device 10 displays the radar chart on the display unit 13 based on the received display data.
[0086] Furthermore, the determination device 30 may include a learning control unit 233 of the learning device 20. In step S204, the learning control unit 233 of the determination device 30 receives information on the correct evaluation score input from the terminal device 10 via the input unit 12, and learns (updates) the determination model using the received evaluation score information and speech feature information as training data.
[0087] The determination system 100 configured as described above uses the audio data information of the episode talk to more accurately determine whether the speaker's speaking style of the episode talk is interesting. Specifically, the determination system 100 executes a learning process using a plurality of pieces of utterance feature information generated using the audio data of a plurality of episode talks and an evaluation score for the interestingness of the speaking style, thereby obtaining a trained model. The determination system 100 then acquires audio data of a new episode talk to be determined, and executes preprocessing based on the acquired data to obtain utterance feature information. The determination system 100 inputs the utterance feature information acquired from the episode talk to be determined into the trained model to obtain an evaluation value.
[0088] Furthermore, the above-mentioned speech feature information correlates with the evaluation value of the entertainingness of a speaking style. For example, entertaining and uninteresting speaking styles have the above-mentioned tendencies. Therefore, it is thought that characteristics appear in the speech feature information depending on the characteristics of the speaking style, resulting in a high correlation with the evaluation value. For example, entertaining speaking styles tend to have longer pauses overall, with greater overall variation, and the length of time between punch lines tends to be longer than the length of time between introduction lines. In accordance with these characteristics, it is thought that features such as the average length of time between punch lines (feature parameter P1), the average length of time between introduction lines (feature parameter P4), and the average and variance of the overall pause length (feature parameters P11 and P12) appear in the speech feature information, resulting in a higher correlation.
[0089] Furthermore, an interesting style of speaking tends to have fewer pronunciations in each utterance section overall, with greater variance, and to have more pronunciations in the punchline than in the introduction. It is also believed that the more pronunciations there are in an utterance section, the longer the duration of that utterance section. In response to these characteristics, features such as the average number of pronunciations in the punchline section (feature parameter P2), the average number of pronunciations in the introduction section (feature parameter P5), the average duration of the punchline section (feature parameter P7), the average duration of the introduction section (feature parameter P8), the average and variance of the total number of pronunciations (feature parameters P13 and P14), and the average and variance of the duration of the total utterance section (feature parameters P17 and P18) are likely to appear in the speech feature information, resulting in higher correlations.
[0090] Furthermore, funny speaking styles tend to have a slower overall speaking speed and a faster speaking speed in the punch line than in the introduction. In response to these characteristics, features such as the speaking speed in the punch line (feature parameter P3), the speaking speed in the introduction (feature parameter P6), and the average and variance of the overall speaking speed (feature parameters P15 and P16) appear in the speech feature information, which is thought to result in a higher correlation.
[0091] Furthermore, since the speech rate is calculated by dividing the number of pronunciations by the duration, feature parameter P7 indicating the average duration of the speech section of the punch line and feature parameter P8 indicating the average duration of the speech section of the introduction can be used as feature parameters P3 and P6 instead of the speech rate of the punch line and the introduction. Furthermore, since the number of pronunciations used to calculate the speech rate is used as feature parameters P2 and P5, feature parameters P3 and P6 do not need to be used.
[0092] Due to the above tendency, by determining the evaluation score of an interesting speaking style using utterance feature amount information, it is possible to more accurately determine how interesting a speaking style is.
[0093] FIG. 13 is a diagram illustrating an example of the hardware configuration of an information processing device 90 applied to this embodiment. The information processing device 90 includes a processor 91, a main memory device 92, a communication interface 93, an auxiliary memory device 94, an input / output interface 95, and an internal bus 96. The processor 91, the main memory device 92, the communication interface 93, the auxiliary memory device 94, and the input / output interface 95 are communicably connected to each other via the internal bus 96. The information processing device 90 may be applied to, for example, the learning device 20 and the determination device 30. In this case, for example, the communication units 21 and 31 may be configured using the communication interface 93. For example, the memory units 22 and 32 may be configured using the auxiliary memory device 94. Furthermore, the control units 23 and 33 may be configured using the processor 91 and the main memory device 92.
[0094] (Variation) In the present embodiment, the terminal device 10 and the determination device 30 are configured as separate devices, but they may also be configured as an integrated device. FIG. 14 is a diagram showing a modified example of the determination device 30 configured in this manner. The determination device 30 shown in FIG. 14 further includes an input unit 34, a display unit 35, and an audio output unit 36. The input unit 34, the display unit 35, and the audio output unit 36 of the determination device 30 shown in FIG. 14 function similarly to the input unit 12, the display unit 13, and the audio output unit 14 of the terminal device 10, respectively. The control unit 33 operates in response to an operation on the input unit 34, performs a determination process using audio data or video data of the episode talk input to the input unit 34, outputs an evaluation score of the determination result using the display unit 35 and the audio output unit 36, and displays a radar chart of the utterance feature information on the display unit 35.
[0095] In this embodiment, the learning device 20 and the determination device 30 are configured as separate devices, but they may also be configured as an integrated device. FIG. 15 is a diagram showing a modified example of the determination device 30 configured in this manner. The storage unit 32 of the determination device 30 shown in FIG. 15 also functions as a teacher data storage unit 322 and a preprocessed teacher data storage unit 323. The teacher data storage unit 322 and the preprocessed teacher data storage unit 323 function similarly to the teacher data storage unit 221 and the preprocessed teacher data storage unit 222 of the learning device 20, respectively. The determination model storage unit 321 also functions as the trained model storage unit 223 of the learning device 20. The control unit 33 of the determination device 30 shown in FIG. 15 also functions as a learning control unit 334. The preprocessing control unit 332 performs not only preprocessing of the determination device 30 (preprocessing of the audio data of the episode talk to be determined) but also preprocessing of the learning device 20 (preprocessing of the audio data of the episode talk of the teacher data). The learning control unit 334 functions in the same manner as the learning control unit 233 of the learning device 20 .
[0096] The learning device 20 may be implemented using a plurality of information processing devices. For example, the learning device 20 may be implemented using a device such as a cloud. For example, in the learning device 20, the memory unit 22 and the control unit 23 may each be implemented in different information processing devices. For example, the memory unit 22 of the learning device 20 may be distributed and implemented across a plurality of information processing devices. The determination device 30 may be implemented using a plurality of information processing devices. For example, the determination device 30 may be implemented using a device such as a cloud. For example, in the determination device 30, the memory unit 32 and the control unit 33 may each be implemented in different information processing devices. For example, the memory unit 32 of the determination device 30 may be distributed and implemented across a plurality of information processing devices.
[0097] An experiment using the determination system 100 of this embodiment will be described. In the experiment, verification was performed using cross-validation (jackknife method). In the verification, 10 divisions were performed 10 times, with 90% training data and 10% verification data. The gradient boosting machine (GBM) was used as the machine learning method. The number of data was 80, with 40 data items being interesting stories and 40 data items being uninteresting stories. The prediction target was a binary value: whether the speaking style was interesting or uninteresting. The following four feature patterns were used. Feature pattern A used feature parameters P1, P2, P4, and P5. Feature pattern B used feature parameters P1 to P6. Feature pattern C used feature parameters P1 to P8. Feature pattern D used feature parameters P11 to P18. The value of feature parameter P3 was calculated by averaging the speaking rates calculated for each speech section in the punch line. Similarly, the value of the feature parameter P6 was calculated by averaging the speech rates calculated for each speech section in the introduction part.
[0098] FIG. 16 shows experimental results when feature pattern A was used, FIG. 17 shows experimental results when feature pattern B was used, FIG. 18 shows experimental results when feature pattern C was used, and FIG. 19 shows experimental results when feature pattern D was used. FIGS. 16 to 18 show prediction accuracy, AUC (Area Under the ROC Curve), and the structure of a prediction model generated by gradient boosting. The weights of feature parameters indicate the importance of the feature parameters. The experimental results shown in FIGS. 16 to 18 reveal that for feature patterns A to C, which use feature parameters obtained from the introductory and punchline sections, respectively, the greater the number of feature parameters, the higher the prediction accuracy. However, even for feature pattern A, which has the fewest number of feature parameters, high prediction accuracy was achieved. It also reveals that the average time length between the introductory sections (feature parameter P4) is the most important. Furthermore, as shown in the experimental results shown in FIG. 19, very high prediction accuracy was also achieved for feature pattern D, which uses feature parameters obtained without separating the introductory and punchline sections. It can be seen that in feature pattern D, the average time length between all of the features (feature parameter P11) is the most important.
[0099] From the above results, it is believed that it is possible to determine whether a speaker's speaking style is interesting by using at least the most important feature parameter, i.e., the average duration between the introductory sections (feature parameter P4) or the average duration between the entire introductory section and punchline section (feature parameter P11), as speech feature information to determine the speaker's speaking style. Furthermore, it is believed that it is possible to more accurately determine whether a speaker's speaking style is interesting by using the next most important feature parameter, for example, the variance of the overall speaking rate (feature parameter P16), as speech feature information. Because the variance of the overall speaking rate is highly important, it is also possible to use the variance of the speaking rate of multiple speech sections belonging to the introductory section as speech feature information.
[0100] According to the above-described embodiment, the determination system includes a learning unit and a determination unit. The determination system may include a learning device and a determination device, and the learning device may include the learning unit and the determination device may include the determination unit. The learning unit uses training data including feature information representing characteristics of the speaker's speaking style obtained based on the duration and number of utterances of each of multiple utterance sections belonging to the punchline part of the speaker's speech, the duration of each of multiple pauses belonging to the punchline part, the duration and number of utterances of each of multiple utterance sections belonging to the introduction part of the speech, and the duration of each of multiple pauses belonging to the introduction part, as well as an evaluation score representing the entertainingness of the speaker's speaking style, to train a determination model that inputs the feature information obtained from the speech of the subject to be evaluated and outputs an evaluation score for the entertaining speaking style of the subject to be evaluated. The determination unit uses the determination model trained by the learning unit to obtain an evaluation score for the entertaining speaking style of the subject to be evaluated corresponding to the feature information obtained from the speech of the subject to be evaluated.
[0101] The feature information includes information on a first feature, and may further include information on one or both of a second feature and a third feature. The first feature includes an average duration of multiple pauses belonging to the punchline, an average number of pronunciations in multiple speech sections belonging to the punchline, an average duration of multiple pauses belonging to the introduction section, and an average number of pronunciations in multiple speech sections belonging to the introduction section. The second feature includes a speech rate of the punchline and a speech rate of the introduction section. The third feature includes an average duration of multiple speech sections belonging to the punchline and an average duration of multiple speech sections belonging to the introduction section.
[0102] Alternatively, the feature information may include the average and variance of the duration of multiple pauses in the entire story, including the punchline and introduction, the average and variance of the number of pronunciations in multiple speech sections in the entire story, the average and variance of the speaking rate in multiple speech sections in the entire story, and the average and variance of the duration of multiple speech sections in the entire story.
[0103] Alternatively, the feature information may include at least an average of the durations of multiple pauses in the entire speech, or an average of the durations of multiple pauses in the introductory part of the speech.Furthermore, the feature information may include a variance of the speech rate of multiple speech sections in the entire speech, or a variance of the speech rate of multiple speech sections in the introductory part of the speech.
[0104] The evaluation system may further include a preprocessing control unit. For example, the evaluation device may include the preprocessing control unit. The preprocessing control unit detects multiple pauses and multiple speech intervals based on the voice volume indicated by the audio data of the speech of the person to be evaluated, defines a predetermined number of the last speech intervals and pauses among the detected multiple speech intervals as punch lines, and defines the section from the start of the speech to the punch line as an introduction line, and acquires feature information based on the pauses and speech intervals included in each of the detected punch lines and introduction lines and the number of pronunciations in each speech interval obtained based on the speech recognition result of the audio data.
[0105] The trained model may be used as a program module that is part of artificial intelligence software. A processor of a computer used as the determination device 30 (e.g., the processor 91 of the information processing device 90 of the embodiment) realizes the function of the determination unit 333 in accordance with instructions from the trained model received from the learning device 20 and stored in a memory (e.g., the main storage device 92 or the auxiliary storage device 94 of the embodiment).
[0106] Although an embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to this embodiment, and includes designs within the scope of the gist of the present invention. [Industrial Applicability]
[0107] It can be used for public speaking courses, speech control for robots and automatic announcements, and automatic evaluation of speaking style. [Explanation of symbols]
[0108] 100...Determination system, 10...Terminal device, 11...Communication unit, 12...Input unit, 13...Display unit, 14...Audio output unit, 15...Memory unit, 16...Control unit, 20...Learning device, 21...Communication unit, 22...Memory unit, 221...Teacher data memory unit, 222...Preprocessed teacher data memory unit, 223...Trained model memory unit, 23...Control unit, 231...Information control unit, 232...Preprocessing control unit, 233...Learning control unit, 24...Input unit, 25...Display unit, 26...Audio output unit, 30...Determination device, 31...Communication unit, 32...Memory unit, 321...Determination model memory unit, 322...Teacher data memory unit, 223...Trained model memory unit, 33...Control unit, 34...Input unit, 35...Display unit, 36...Audio output unit, 331...information control unit, 332...preprocessing control unit, 333...determination unit, 334...learning control unit, 91...processor, 92...main memory device, 93...communication interface, 94...auxiliary memory device, 95...input / output interface, 96...internal bus
Claims
1. a learning unit that uses training data including feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introduction portion of the speech, and the duration of each of a plurality of pauses belonging to the introduction portion, and an evaluation score representing the entertainingness of the speaking style of the speaker, as input of the feature information obtained from the speech of a subject to be evaluated, and learns a judgment model that outputs an evaluation score for the entertainingness of the speaking style of the subject to be evaluated; A learning device comprising:
2. The feature information is a first feature amount including an average of a plurality of pause durations belonging to the punchline portion, an average of the number of pronunciations of a plurality of speech sections belonging to the punchline portion, an average of a plurality of pause durations belonging to the introduction portion, and an average of the number of pronunciations of a plurality of speech sections belonging to the introduction portion; a second feature amount including a speech rate of the punch line portion and a speech rate of the introduction portion; and a third feature including an average of the durations of a plurality of speech sections belonging to the punch line portion and an average of the durations of a plurality of speech sections belonging to the introduction portion, The learning device according to claim 1 .
3. a judgment unit that receives feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and that obtains an evaluation score for the entertainingness of the speaking style of the speaker corresponding to the feature information obtained from the speech of the person to be evaluated, using a judgment model that outputs an evaluation score representing the entertainingness of the speaking style of the speaker; A determination device comprising:
4. The feature information is a first feature amount including an average of a plurality of pause durations belonging to the punchline portion, an average of the number of pronunciations of a plurality of speech sections belonging to the punchline portion, an average of a plurality of pause durations belonging to the introduction portion, and an average of the number of pronunciations of a plurality of speech sections belonging to the introduction portion; a second feature amount including a speech rate of the punch line portion and a speech rate of the introduction portion; and a third feature including an average of the durations of a plurality of speech sections belonging to the punch line portion and an average of the durations of a plurality of speech sections belonging to the introduction portion, The determination device according to claim 3 .
5. a pre-processing control unit that detects a plurality of pauses and a plurality of speech sections based on the volume of a voice indicated by the speech data of the person to be evaluated, defines a predetermined number of the last speech sections and pauses among the detected plurality of speech sections as punch lines, and defines a section from the start of the speech to the punch line as an introduction line, and acquires the feature information based on the pauses and speech sections included in the detected punch lines and introduction lines, respectively, and the number of pronunciations in each of the speech sections obtained based on a speech recognition result of the speech data. The determination device according to claim 3 .
6. a learning unit that uses training data including feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and an evaluation score representing the entertainingness of the speaking style of the speaker, as input of the feature information obtained from the speech of a subject to be evaluated, and learns a judgment model that outputs an evaluation score for the entertainingness of the speaking style of the subject to be evaluated; a judgment unit that uses the trained judgment model to obtain an evaluation score corresponding to the feature amount information obtained from the speech of the person to be evaluated; A determination system comprising:
7. a learning step of using training data including feature information representing characteristics of the speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of the speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and an evaluation score representing the entertainingness of the speaking style of the speaker, as input of the feature information obtained from the speech of the person to be evaluated, and training a judgment model that outputs an evaluation score for the entertainingness of the speaking style of the person to be evaluated; A learning method that has
8. a judgment step of inputting feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line portion of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line portion, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory portion of the speech, and the duration of each of a plurality of pauses belonging to the introductory portion, and using a judgment model that outputs an evaluation score representing the entertainingness of the speaking style of the speaker, and obtaining an evaluation score for the entertainingness of the speaking style of the person to be evaluated corresponding to the feature information obtained from the speech of the person to be evaluated; A determination method having the following.
9. a trained model trained using training data including feature information representing characteristics of a speaking style obtained based on the duration and number of pronunciations of each of a plurality of speech sections belonging to a punch line part of a speaker's speech, the duration of each of a plurality of pauses belonging to the punch line part, the duration and number of pronunciations of each of a plurality of speech sections belonging to an introductory part of the speech, and the duration of each of a plurality of pauses belonging to the introductory part, and an evaluation score representing the entertainingness of the speaking style of the speaker, A trained model for inputting the feature information obtained from the speech of a person to be evaluated into a computer and for outputting an evaluation score for the person's interesting speaking style.
10. Computer, A program for causing the learning device according to claim 1 or 2 to function.
11. Computer, A program for causing the determination device according to any one of claims 3 to 5 to function.
12. a learning step of using training data including feature information of a speaker's speaking style, which is an average of multiple pause lengths in the speaker's speech, and an evaluation score representing the interestingness of the speaker's speaking style, to learn a judgment model that inputs the feature information obtained from the speech of a subject to be evaluated and outputs an evaluation score for the interestingness of the speaking style of the subject to be evaluated; A learning model generation method having the following.
13. a learning step of learning a judgment model that inputs the feature information obtained from the speech of a subject to be evaluated and outputs an evaluation score for the interestingness of the speech of the subject to be evaluated, using training data including an average of multiple pause lengths belonging to the introductory part of the speech of the subject, which is feature information of the speech of the subject, and an evaluation score representing the interestingness of the speech of the subject to be evaluated; A learning model generation method having the following.
14. the feature information further includes a variance of a speaking rate of a plurality of speech sections in the speech of the speaker; The learning model generation method according to claim 12 or 13.
15. the feature information further includes a variance of the speaking rate of a plurality of speech sections belonging to an introductory part of the speech of the speaker; The learning model generation method according to claim 12 or 13.
Citation Information
Patent Citations
Amusingness quantitative evaluation device, amusingness quantitative evaluation method, and program
JP2018022118A