Voice analysis system and voice analysis method
The speech analysis system improves voice emotion recognition accuracy in noisy environments by using distinct thresholds for speech and emotion recognition, effectively filtering noise and enhancing emotion detection precision.
Patent Information
- Application Number
- JP2024037433
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-11
- Publication Date
- 2025-09-25
AI Technical Summary
Conventional speech recognition systems in noisy environments suffer from reduced accuracy due to the inclusion of noise in speech intervals, which affects the precision of voice emotion recognition.
A speech analysis system that employs separate threshold settings for speech recognition and voice emotion recognition, using a higher threshold for voice emotion recognition to exclude noise and improve accuracy by detecting speech intervals specifically for emotion recognition.
Enhances the accuracy of voice emotion recognition in noisy environments by effectively distinguishing speech segments from noise, allowing for more precise emotion detection and analysis.
Smart Images

Figure 2025138379000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a speech analysis system and a speech analysis method. [Background technology]
[0002] Conventionally, various measures have been taken to objectively evaluate and improve the quality of services involving dialogue between customers and operators, such as contact centers and front desk services. For this purpose, speech recognition techniques, which recognize a speaker's utterances from their voice, have come into use. Using speech recognition techniques, for example, it is possible to transcribe speech from speech into text and obtain the speech emotion corresponding to the text. For this reason, there is a growing demand for utilizing the text of speech and the speech emotion.
[0003] In contact centers and other front-line work, noise is likely to occur due to nearby operators and air conditioning operating indoors. In such noisy environments, the accuracy with which a computer recognizes vocal emotion deteriorates. Therefore, there is a need for highly accurate vocal emotion recognition even in noisy environments.
[0004] A known technique for outputting the emotion contained in the speech along with the speech recognition result is the technology disclosed in Patent Document 1. Patent Document 1 describes the system as "a speech emotion recognition system comprising a speech analysis unit that performs speech analysis processing on captured speech, an acoustic model unit that has speech patterns in phonemes, a vocal transformation emotion model unit that represents transformations of the phoneme spectrum due to emotion, and a speech recognition unit that performs speech recognition processing on the speech analysis result by linking the acoustic model unit, the vocal transformation emotion model unit, and a dictionary unit, and that outputs words and sentences to be recognized as speech recognition results based on the speech features, and also outputs a level that indicates the degree of emotion of the speaker that the speech conveys." [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 11-119791 Summary of the Invention [Problem to be solved by the invention]
[0006] One embodiment of the invention disclosed in Patent Document 1 applies speech recognition and speech emotion recognition to speech intervals obtained by performing speech interval detection on recorded audio data. An utterance interval is, for example, a time interval from the start to the end of a speaker's speech. Furthermore, a speech interval divided into smaller parts is called a speech segment. A speech segment corresponds to a specific sentence, phrase, or word in the speech interval. Conventional speech recognition processing will now be described with reference to FIG. 1.
[0007] 1 is a diagram showing an example of a speech section detected from a speech, in which the speech section detected from the speech differs depending on a speech section threshold such as sound pressure. (1) The upper part of Figure 1 shows an example of a speech segment detected from audio when the speech segment threshold is set low. In this example, noise before and after the actual speech segment "Payment is not going well" is also detected as a speech segment.
[0008] (2) The bottom of Figure 1 shows an example of a speech interval detected from audio when the speech interval threshold is increased. In this example, noise before and after the actual speech is not included, but only a portion of the speech is detected as a speech interval.
[0009] In general, the sound pressure is low at the beginning and end of a speech. Conventionally, to ensure the accuracy of speech recognition, a wide speech interval is detected by setting a low speech interval threshold. However, in a noisy environment, the speech interval detected by setting a low speech interval threshold includes not only the speech segment but also noise before and after the speech segment. Using a speech interval that includes noise in this way reduces the accuracy of voice emotion recognition.
[0010] The present invention has been made in view of the above circumstances, and has an object to improve the accuracy of voice emotion recognition. [Means for solving the problem]
[0011] The speech analysis system of the present invention comprises a speech input unit that inputs a speech signal, a speech interval detection unit for speech recognition that detects a first speech interval for speech recognition from the speech signal that exceeds a speech interval threshold for detecting a first speech interval for speech recognition and extracts the first speech segment, a speech interval detection unit for speech emotion recognition that detects a second speech interval for speech emotion recognition from the speech signal that exceeds a speech emotion recognition threshold that is higher than the speech interval threshold and extracts the second speech segment, a speech emotion recognition unit that performs speech emotion recognition processing based on the second speech segment extracted from the second speech interval and assigns emotion to the second speech segment, and an output unit that associates emotion with the textualized speech and outputs the textualized speech and emotion. [Effects of the Invention]
[0012] According to the present invention, it is possible to improve the accuracy of voice emotion recognition. Problems, configurations, and effects other than those described above will become apparent from the following description of the embodiments. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 10 is a diagram illustrating an example of a speech section detected from a voice. [Figure 2] 1 is a hardware configuration diagram of a voice analysis system according to a first embodiment of the present invention. [Figure 3] 2 is a block diagram showing an example of the functional configuration of a voice analysis execution unit according to the first embodiment of the present invention. FIG. [Figure 4] FIG. 2 is a diagram showing an example of detecting a speech segment by changing a speech segment threshold based on a voice (voice signal) input to the voice input device according to the first embodiment of the present invention. [Figure 5] 5 is a flowchart showing an example of processing by a voice analysis execution unit according to the first embodiment of the present invention. [Figure 6]FIG. 10 is a block diagram showing an example of the configuration of a voice analysis execution unit having an external voice recognition unit according to a modified example of the first embodiment of the present invention. [Figure 7] FIG. 10 is a hardware configuration diagram of a voice analysis system according to a second embodiment of the present invention. [Figure 8] FIG. 10 is a block diagram showing an example of the functional configuration of a voice analysis execution unit according to the second embodiment of the present invention. [Figure 9] FIG. 10 is a diagram showing how emotion labels are assigned to speech segments according to the second embodiment of the present invention. [Figure 10] 10 is a flowchart showing an example of processing by a voice analysis execution unit according to the second embodiment of the present invention. [Figure 11] FIG. 10 is a block diagram showing an example of the configuration of a voice analysis execution unit that is provided externally in a voice recognition system according to a modified example of the second embodiment of the present invention. [Figure 12] FIG. 2 is an interface diagram showing a first display example of a voice analysis result according to the first and second embodiments of the present invention. [Figure 13] FIG. 10 is an interface diagram showing a second display example of the voice analysis result according to the first and second embodiments of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted.
[0015] [First embodiment] Fig. 2 is a hardware configuration diagram of a voice analysis system 1 according to a first embodiment of the present invention. As shown in Fig. 2, the voice analysis system 1 is configured with an information processing device such as a server. The voice analysis system 1 may be, for example, an on-premise server installed within a company, or a cloud server installed outside the company and accessible online from a PC (Personal Computer). Alternatively, the voice analysis system 1 may be configured as a standalone system using a program installed on a PC.
[0016] The voice analysis system 1 has a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, and a voice input device 14 such as a microphone. All of these components are connected to each other via a bus 10, and perform data input and output between them. The CPU 11 is generally called a processing device, the ROM 12 and RAM 13 form a storage device, and the voice input device 14 is an example of an input device.
[0017] The CPU 11 reads out program code of software that realizes each function according to this embodiment from the ROM 12, loads it into the RAM 13, and executes it. Variables, parameters, etc. generated during the calculation process of the CPU 11 are temporarily written to the RAM 13, and these variables, parameters, etc. are read out by the CPU 11 as appropriate. The ROM 12 records programs, data, etc. necessary for the operation of the CPU 11. In other words, the ROM 12 is used as an example of a computer-readable non-transitory storage medium that stores programs executed by the CPU 11.
[0018] A voice analysis execution unit 20 shown in Fig. 3, which will be described later, is implemented in the RAM 13 by the program code of the software read from the ROM 12 and loaded. The voice analysis execution unit 20 has a function of processing input from the voice input device 14 and performing voice analysis. The voice analysis execution unit 20 will be described in detail below.
[0019] FIG. 3 is a block diagram showing an example of the functional configuration of the voice analysis execution unit 20 according to this embodiment. As shown in FIG. 3, the voice analysis execution unit 20 includes a voice input unit 21, a voice recognition speech period detection unit 22, a voice recognition unit 23, a voice emotion recognition speech period detection unit 24, a voice emotion recognition unit 25, and an output unit 26.
[0020] The voice input unit 21 converts analog voice input from the voice input device 14 into a digital voice signal. Then, the voice input unit 21 inputs the voice signal to the voice recognition speech period detector 22 and the voice emotion recognition speech period detector 24.
[0021] The speech recognition speech interval detection unit 22 detects a speech interval for speech recognition from the speech signal exceeding a speech interval threshold for detecting a speech recognition speech interval (an example of a first speech interval) for speech recognition, and extracts an utterance segment (an example of a first speech segment). For example, the speech recognition speech interval detection unit 22 compares the sound pressure of the speech signal input from the speech input unit 21 with the speech interval threshold, and detects a speech interval for speech recognition from the speech signal at a portion where the sound pressure of the speech signal exceeds the speech interval threshold.
[0022] The speech recognition unit 23 performs speech recognition processing based on the speech section for speech recognition, converts the speech into text, and outputs the textual speech to the output unit 26. For example, the speech recognition unit 23 performs speech recognition for each speech segment in the speech section for speech recognition detected by the speech section detection unit for speech recognition 22.
[0023] The speech interval detection unit for voice emotion recognition 24 detects a speech interval for voice emotion recognition (an example of a second speech interval) from the speech signal that exceeds a higher threshold for voice emotion recognition than the speech interval threshold, and extracts a speech segment (an example of a second speech segment). For example, the speech interval detection unit for voice emotion recognition 24 compares the sound pressure of the speech signal input from the speech input unit 21 with the threshold for voice emotion recognition, and detects a speech interval for voice emotion recognition from the speech signal where the sound pressure exceeds the threshold for voice emotion recognition. Here, by setting the threshold for voice emotion recognition, the speech interval detection unit for voice emotion recognition 24 divides the speech interval for voice emotion recognition into utterance units and removes noise mixed in the speech. Therefore, the speech segments extracted by the speech interval detection unit for voice emotion recognition 24 from the speech signal are either the same length as or shorter than the speech segments extracted by the speech interval detection unit for voice recognition 22.
[0024] The voice emotion recognition unit 25 performs voice emotion recognition processing based on the speech segments extracted in the speech sections for voice emotion recognition, and assigns emotions to the speech segments. For example, the voice emotion recognition unit 25 assigns the voice emotions recognized based on the speech segments extracted in the speech sections for voice emotion recognition as emotion labels to the speech segments. The process of assigning emotion labels to speech segments is also referred to as "assigning emotions to speech segments."
[0025] The output unit 26 associates emotions with the converted voice text and outputs the converted voice text and the emotions. The operator can check the converted voice text and emotions output by the output unit 26 and understand the emotions of the other party. The converted voice text is input from the voice recognition unit 23. As shown in FIG. 6, which will be described later, the conversion of voice text may be performed by an external voice recognition system 30.
[0026] Next, speech recognition and speech emotion recognition will be described with reference to FIG. 4 is a diagram showing an example of detecting a speech section by changing the speech section threshold based on the speech (speech signal) input to the speech input device 14. The speech section threshold makes it possible to distinguish between speech and noise when detecting them.
[0027] As shown on the left side of Figure 4, the speech section for speech recognition detected from the speech signal using a low speech section threshold includes not only the actual utterance "Payment is not going well" but also noise before and after this utterance. However, the speech section detection unit for speech recognition 22 performs speech recognition based on the speech section for speech recognition, ignoring the noise and outputting the text version of the utterance "Payment is not going well."
[0028] As shown on the right side of Figure 4, the speech section for voice emotion recognition detected by increasing the speech section threshold includes part of the actual speech. Therefore, the speech section detection unit for voice emotion recognition 24 performs voice emotion recognition based on the speech section for voice emotion recognition, and assigns the emotion of "anger" to the speech segment of this speech section and outputs it.
[0029] 5 is a flowchart showing an example of processing by the voice analysis execution unit 20 according to this embodiment. An example of the operation of the voice analysis execution unit 20 will be described below with reference to FIGS.
[0030] First, the audio input unit 21 converts an analog input audio signal input from a microphone into a digital audio signal (abbreviated as audio signal) using an AD converter. This digital audio signal may be an audio signal that has been subjected to dereverberation, speech enhancement, or sound source separation in advance by another audio system, for example.
[0031] The speech recognition speech activity detector 22 performs speech activity detection based on the audio signal (S1). Speech activity detection, also called voice activity detection or VAD (Voice Activity Detection), is a technique for determining the presence of speech in an audio signal that contains both speech and non-speech sounds. The speech recognition speech activity detector 22 performs speech activity detection by setting a speech activity threshold such as sound pressure. In addition to sound pressure, frequency may also be set as the speech activity threshold, since people's voices tend to get higher when excited.
[0032] As shown in the example of Fig. 1, the speech segments used in speech recognition must include all of the parts that were actually spoken. For this reason, the speech recognition speech period detection unit 22 sets the speech period threshold to a low value to perform speech period detection.
[0033] In step S1, the speech interval detection unit for voice emotion recognition 24 performs speech interval detection based on the same audio signal as the speech interval detection unit for voice recognition 22. Therefore, the speech interval detection unit for voice emotion recognition 24 sets a speech interval recognition threshold that is higher than the speech interval threshold used by the speech recognition speech interval detection unit for voice recognition 22 for speech interval detection, and limits the speech interval to be detected. Therefore, as shown in Fig. 4, the speech interval detection unit for voice emotion recognition 24 can more actively remove noise contained before and after the speech segments used in speech emotion recognition, thereby improving the accuracy of speech emotion recognition in noisy environments.
[0034] It is also possible that the speech recognition speech interval detector 22 and the speech emotion recognition speech interval detector 24 are integrated into a single speech interval detection module using a common speech interval threshold, and then perform speech recognition and speech emotion recognition processing on each obtained speech segment. However, as shown in Figure 4, when a low speech interval threshold is set for speech recognition, noise is likely to be included in the speech interval. For this reason, in a noisy environment, a method using a common threshold cannot perform speech emotion recognition with as much accuracy as the method according to this embodiment, which uses different thresholds (a speech interval threshold and a threshold for speech emotion recognition).
[0035] The speech recognition speech activity detector 22 and the speech emotion recognition speech activity detector 24 may estimate the probability of a speech activity using, but are not limited to, methods such as logistic regression from acoustic features such as fundamental frequency and Mel-Frequency Cepstrum Coefficients (MFCCs) in addition to sound pressure. The speech activity detection threshold may be any value that indicates the probability that a detected segment is a speech activity. In addition to the sound pressure value described above, other acoustic feature values such as the fundamental frequency and Mel-frequency cepstrum coefficients, or speech activity probability values output by methods such as logistic regression may also be used.
[0036] After step S1, the voice recognition unit 23 and the voice emotion recognition unit 25 perform voice recognition and voice emotion recognition based on each speech segment (S2), and then this process ends.
[0037] In step S2, the speech recognition unit 23 determines a text representation of the encoded speech (digital speech signal) based on each speech segment obtained from the speech recognition speech activity detection unit 22, and performs transcription. Here, speech recognition processing may be performed using a method that outputs a text representation from the speech signal by using an acoustic model that identifies phonemes from the speech signal, a pronunciation dictionary in which words and their pronunciations are registered, and a language model that determines combinations of words with a high probability of occurrence and converts them into sentences. A combination of words with a high probability of occurrence refers to a combination of words that is used when determining which word follows a certain word, taking context into consideration. For example, if there is a word such as "payment goes well," there is a high probability that this word will be followed by a word such as "it doesn't go well," so speech recognition will be performed as "payment doesn't go well."
[0038] The speech recognition unit 23 may also perform speech recognition processing using a method such as combining the acoustic model, speech dictionary, and language model into a single neural network to output a text representation from a speech signal. Furthermore, the speech recognition unit 23 is not limited to these methods, and may perform speech recognition processing using various other methods.
[0039] In step S2, the voice emotion recognition unit 25 estimates the emotion contained in the encoded voice (digital voice signal) based on the speech segments obtained from the voice emotion recognition speech period detection unit 24 in step S1. Emotions can be expressed using categories such as "joy" or "anger" or coordinate values in an emotional space with two axes such as "pleasant-unpleasant" or "awake-asleep." As shown in Figure 4, voice emotion recognition differs from speech recognition in that it treats utterances in which the beginning or end of an utterance is missing as containing emotion. This allows the voice emotion recognition unit 25 to estimate the emotion contained in the voice without sacrificing accuracy.
[0040] Furthermore, when estimating voice emotion by the voice emotion recognition unit 25, a method may be used in which a feature vector such as a logarithmic power spectrum or a Mel filter bank is extracted from the voice signal and input to a statistical classifier such as linear discriminant analysis or a support vector machine (SVM), a regression model such as linear regression or support vector regression (SVR), or a neural network. Alternatively, the voice emotion recognition unit 25 may use an end-to-end method in which the voice waveform is input directly to a neural network model without extracting a feature vector. Furthermore, the voice emotion recognition unit 25 is not limited to these methods, and may use various other methods to perform voice emotion recognition processing.
[0041] According to the voice analysis system 1 of the first embodiment described above, when detecting a speech section for voice emotion recognition, the system detects the speech section using a threshold for voice emotion recognition that is higher than the speech section threshold. This makes it possible to perform voice emotion recognition with high accuracy while suppressing the influence of noise. As a result, it becomes possible to detect a speech section that is more appropriate for the purpose than a conventional speech section that is detected using a common threshold for voice recognition and voice emotion recognition.
[0042] Furthermore, the accuracy of speech emotion recognition processing is improved even in noisy environments. As a result, for example, operators can use the emotion labels output by the output unit 26 as appropriate and analyze customer service operations.
[0043] [Modification of the first embodiment] The voice recognition unit 23 of the voice analysis execution unit 20 may be replaced with an external voice recognition system. 6 is a block diagram showing an example of the configuration of a voice analysis execution unit 20A that has an external voice recognition system 30. The voice recognition system 30 performs voice recognition on a cloud server or the like, and provides text data obtained by transcribing the voice.
[0044] The voice analysis execution unit 20A does not include the voice recognition unit 23 shown in Fig. 3. However, the voice analysis execution unit 20A causes an external voice recognition system 30 to perform voice recognition and voice-to-text processing. For this reason, the voice input unit 21 outputs voice data to the voice recognition system 30. The functions of the voice analysis execution unit 20A other than voice recognition are the same as those of the voice analysis execution unit 20 shown in Fig. 3.
[0045] The voice recognition system 30 performs voice recognition based on the voice data input from the voice analysis execution unit 20A and converts the voice into text. The voice recognition system 30 then outputs the text voice data to the voice analysis execution unit 20A. The text voice data is accompanied by text and a timestamp.
[0046] The output unit 26 associates the emotion label of the voice emotion recognized by the voice emotion recognition unit 25 with the text-converted voice data input from the voice recognition system 30 and outputs the text. At this time, the output unit 26 associates the timestamp of the speech section detected by the voice recognition speech section detection unit 22 with the timestamp of the voice data input from the voice recognition system 30. Thereafter, the output unit 26 uses the timestamp to associate the emotion label with the text-converted voice at the correct timing and outputs the text.
[0047] The voice analysis execution unit 20A according to the modification of the first embodiment can reduce the processing load of the voice analysis execution unit 20A itself by having the voice recognition process performed by the voice recognition system 30. Furthermore, the system configuration of the voice analysis execution unit 20A can be simplified.
[0048] [Second embodiment] Next, a voice analysis system according to a second embodiment of the present invention will be described with reference to FIG. 7 and subsequent figures. Fig. 7 is a hardware configuration diagram of a speech analysis system 1A according to the second embodiment. Components similar to those in Fig. 2 are given the same reference numerals and descriptions thereof will be omitted. An output device 15 configured by a display or the like is added to the speech analysis system 1A. All of the components are connected to each other via a bus 10 or the like, and perform input and output of data between each other.
[0049] A voice analysis execution unit 20B shown in FIG. 8, which will be described later, is implemented in the RAM 13 by a software program. The voice analysis execution unit 20B processes the voice input from the voice input device 14, performs voice analysis, and outputs the integrated result. The output device 15 displays the output result. The voice analysis execution unit 20B and the output device 15 will be described in detail below.
[0050] 8 is a block diagram showing an example of the functional configuration of the voice analysis execution unit 20B according to this embodiment. The same components as those in FIG. 3 are given the same reference numerals and their description will be omitted.
[0051] The voice analysis execution unit 20B includes an emotion selection unit 27 in addition to the functional units included in the voice analysis execution unit 20 shown in FIG.
[0052] The emotion selection unit 27 selects an emotion to be output to the output unit 26 based on the emotion assigned to the speech segment extracted in the speech section for speech emotion recognition by the voice emotion recognition unit 25. At this time, the emotion selection unit 27 integrates the output from the voice recognition unit 23 and the output from the voice emotion recognition unit 25. For example, if the voice emotion recognition unit 25 recognizes multiple voice emotions for one speech segment extracted in the speech section for speech recognition, the emotion selection unit 27 performs processing to select one voice emotion from the multiple voice emotions.
[0053] If there is no result of the voice emotion recognition processing for the utterance period for voice recognition, the emotion selection unit 27 changes the speech period threshold to a value lower than the previous value and performs feedback processing to have the utterance period detection unit for voice recognition 22 redetect the utterance period for voice recognition. In this case, the emotion selection unit 27 may instruct the utterance period detection unit for voice emotion recognition 24 to perform reprocessing.
[0054] The output unit 26 outputs the final result integrated by the emotion selection unit 27. The output unit 26 displays the result on the screen of the output device 15, for example, so that the user can confirm the result.
[0055] Here, emotion labels will be described with reference to FIG. Fig. 9 is a diagram showing how emotion labels are assigned to speech segments. Fig. 9 shows examples A to C of speech intervals, each of which is shaded with diagonal lines. The speech interval in example A is short, while the speech intervals in examples B and C are long. Additionally, noise is included at the end of the speech interval in example B.
[0056] (1) Speech segments for speech recognition Speech segments that can be recognized by voice are detected in all of the speech sections of Examples A to C. However, the lengths of the speech sections detected in Examples A to C and the speech segments that can be recognized by voice are different.
[0057] (2) Speech segments for voice emotion recognition No speech segments that allow for speech emotion recognition were extracted from the speech section of Example A. Two speech segments that allow for speech emotion recognition were extracted from the speech section of Example B. Furthermore, one speech segment that allows for speech emotion recognition was extracted from the speech section of Example C. Therefore, an emotion label that matches the recognized speech emotion is assigned.
[0058] (3) Output Since no voice emotion has been recognized in the speech segment for voice emotion recognition in Example A, the emotion label "Neutral" is output from the output unit 26. Note that, for a speech segment in which no voice emotion has been recognized, such as Example A, no emotion label may be output from the output unit 26.
[0059] In the speech segment for voice emotion recognition of example B, emotion labels of "anger" and "sadness" are output from the output unit 26 as two recognized voice emotions. Note that the two recognized voice emotions may be the same. Also, only the emotion label of "anger", which is the longer of the speech segment for voice emotion recognition, may be output from the output unit 26. In the speech segment for voice emotion recognition of example C, emotion label of "anger" is output from the output unit 26 as one recognized voice emotion.
[0060] Fig. 10 is a flowchart showing an example of processing by the voice analysis execution unit 20B according to this embodiment. An example of the operation of the voice analysis execution unit 20B will be described below with reference to Fig. 8 and Fig. 10. Note that the processing in steps S1 and S2 is the same as the processing in each step in Fig. 5, and therefore detailed description thereof will be omitted.
[0061] After step S2, the speech recognition result output from the speech recognition unit 23 and the speech emotion result output from the speech emotion recognition unit 25 are input to the emotion selection unit 27, which performs processing. If the speech recognition unit 23 and the speech emotion recognition unit 25 output results for speech segments in different speech sections, the emotion selection unit 27 assigns emotions by performing various conditional branching in the processing from step S3 to step S10, and integrates the selected emotions.
[0062] As explained in the first embodiment, the speech activity detection unit for voice emotion recognition 24 sets a higher speech activity threshold than the speech activity threshold used by the speech activity detection unit for voice recognition 22, and actively removes noise before and after the speech segment. For example, as shown in FIG. 4, in a noisy environment, speech segments for voice emotion recognition are likely to be output with a shorter, more limited length than speech segments for speech recognition. The processing in this case will be explained in detail below. Note that the processing and conditional branching from step S3 to step S10 are performed for each speech segment obtained by the speech activity detection for voice recognition.
[0063] The emotion selection unit 27 determines whether or not a speech recognition result exists for the speech recognition utterance section obtained in step S1 (S3). If the emotion selection unit 27 determines that a speech recognition result does not exist for the speech recognition utterance section (NO in S3) and no speech recognition result can be obtained, the emotion selection unit 27 does not output any data for that speech section and ends this process. For example, if the sound in the detected speech recognition utterance section is unclear or contains only noise, no speech recognition result exists and the process ends. Note that the emotion selection unit 27 can also output speech sections for which no speech recognition result exists as noise, etc.
[0064] In step S3, if the emotion selecting unit 27 determines that a speech recognition result exists for the speech recognition utterance section (YES in S3), the process proceeds to conditional branch S4.
[0065] Next, the emotion selection unit 27 determines whether the lengths of the utterance interval for speech recognition and the utterance interval for speech emotion recognition are different (S4). If the emotion selection unit 27 determines that the lengths of the utterance interval for speech recognition and the utterance interval for speech emotion recognition are the same (NO in S4), the emotion selection unit 27 proceeds to step S10, which is connected to connector A. In step S10, if the utterance interval for speech recognition and the utterance interval for speech emotion recognition are the same, the emotion selection unit 27 performs a process of assigning one emotion label to the utterance interval for speech emotion recognition.
[0066] Since the speech segments for speech recognition and speech segments for speech emotion recognition have the same length, there is only one speech emotion recognition result obtained by the speech emotion recognition process for the speech segments for speech emotion recognition. Therefore, the emotion selection unit 27 assigns one emotion label obtained in step S2 to all of the speech recognition results (S10), and proceeds to step S11.
[0067] In step S4, if the emotion selecting unit 27 determines that the lengths of the utterance segments for voice recognition and the utterance segments for voice emotion recognition are different (YES in S4), the process proceeds to conditional branch S5.
[0068] Following the processing of step S4, the emotion selection unit 27 determines whether or not a voice emotion recognition result exists for the speech section for voice recognition (S5). For example, if a voice emotion is recognized from a speech segment for voice emotion recognition, that is, if a speech segment for voice emotion recognition exists, a voice emotion recognition result exists. On the other hand, if the speech segment for voice emotion recognition contains noise or the like and the voice emotion is not recognized, that is, if a speech segment for voice emotion recognition does not exist, no voice emotion recognition result exists.
[0069] If the emotion selection unit 27 determines in step S5 that there is no speech segment for voice emotion recognition (NO in S5), an emotion label cannot be obtained. Therefore, the emotion selection unit 27 changes the voice emotion recognition threshold so that there is a speech segment for voice emotion recognition (S6).
[0070] If the process proceeds to step S6 after the NO determination in step S5, the emotion selection unit 27 changes the voice emotion recognition threshold to a value higher than the previous value (S6). Then, the process proceeds to step S1, where the voice emotion recognition speech period detection unit 24 again detects a voice emotion recognition speech period using the changed voice emotion recognition threshold (S1) and extracts speech segments for voice emotion recognition. Next, the voice emotion recognition unit 25 performs voice emotion recognition on the speech segments of the voice emotion recognition speech period (S2). The processes from step S3 onwards are as described above.
[0071] Note that after a NO determination in step S5, the process may not proceed to step S6. In this case, the output unit 26 may output the image without assigning an emotion label, as in example A shown in Fig. 9 (S11). Furthermore, if the emotion selection unit 27 does not recognize an assignable emotion, the emotion selection unit 27 may perform processing such as assigning an emotion label such as "neutral," and the output unit 26 may output that emotion label.
[0072] In step S5, if the emotion selection unit 27 determines that a voice emotion recognition segment exists in the speech recognition utterance section (YES in S5), an emotion label is obtained, and the process proceeds to conditional branch S7.
[0073] Next, the emotion selection unit 27 determines whether the confidence level of the voice emotion recognition result is higher than a confidence level threshold (S7). The confidence level can be determined using the length of the speech interval for voice emotion recognition, the probability value of the emotion label output by the voice emotion recognition unit 25, or the like. The probability value of the emotion label is, for example, 75% for the voice emotion "anger" and 25% for "impatience." However, the probability value of the emotion label may be calculated as a numerical value between "0" and "1.0" instead of a percentage. Here, if an ambiguous result is obtained, such as "0.4" for the voice emotion "anger" and "0.4" for "neutral," it may be unclear which emotion it represents, and the result may be unreliable. In such a case, the confidence level is determined to be "0.2," which is less than the confidence level threshold of "0.5," and step S7 is determined to be NO.
[0074] Therefore, if there is a result of voice emotion recognition processing for the speech segment for voice recognition (YES in S5) and the confidence level of the result of the voice emotion recognition processing is equal to or lower than the confidence level threshold (NO in S7), the emotion selection unit 27 changes the voice emotion recognition threshold to a value lower than the previous value (S6). By using the changed voice emotion recognition threshold, the voice emotion recognition speech segment detection unit 24 can detect a more appropriate voice emotion recognition speech segment than in the previous processing and extract a speech segment for voice emotion recognition. Therefore, the emotion selection unit 27 causes the voice emotion recognition speech segment detection unit 24 to redetect the voice emotion recognition speech segment.
[0075] After step S6, the process proceeds to step S1, where the speech interval detection unit for voice emotion recognition 24 again detects a speech interval for voice emotion recognition using the changed threshold for voice emotion recognition (S1) and extracts speech segments for voice emotion recognition. The speech emotion recognition unit 25 performs speech emotion recognition on the obtained speech segments (S2). Next, the voice emotion recognition unit 25 performs voice emotion recognition on the speech segments in the speech interval for voice emotion recognition (S2). The processing from step S3 onwards is as described above.
[0076] In step S7, if the confidence level is higher than the confidence level threshold (YES in S7), the process proceeds to conditional branch S8.
[0077] If the confidence level does not improve even after repeated processing, the output unit 26 may output the recognized emotion even if the emotion is incorrect. Therefore, even if the confidence level does not improve after the NO determination in step S7 and the repeated processing of steps S6, S1, S2 and subsequent steps, the process proceeds to conditional branch S8. The number of repeated processing times at which it is determined that the confidence level does not improve is about two or three times.
[0078] Next, the emotion selection unit 27 determines whether or not multiple voice emotion recognition results exist for the speech section for voice recognition (S8). If the emotion selection unit 27 determines that multiple voice emotion recognition results exist (YES in S8), it can assign one of the multiple emotions to one speech segment included in the speech section for voice emotion recognition.
[0079] For this reason, when there are multiple voice emotion recognition processing results for an utterance section for voice recognition, the emotion selection unit 27 integrates the multiple emotion labels and assigns them to the utterance section for voice recognition (S9). This process of assigning one emotion label selected from multiple emotion labels to one speech segment included in an utterance section for voice emotion recognition is called "integrating multiple emotion labels."
[0080] For example, as shown in Example B in Figure 9, the emotion selection unit 27 divides the speech interval for voice emotion recognition in the middle, and integrates multiple emotion labels to assign them to a single speech interval for voice recognition. Note that various processes are conceivable for the emotion selection unit 27 to assign emotion labels to the speech interval for voice recognition in step S9. For example, the emotion selection unit 27 may select and assign the emotion label with the longest speech interval, the emotion label with the largest number of emotion labels, or the emotion label with the largest total probability value output by the voice emotion recognition unit 25, but is not limited to these.
[0081] On the other hand, if it is determined that there are not multiple voice emotion recognition results for the speech section for voice recognition (NO in S8), the emotion selection unit 27 assigns one emotion label to the entire speech recognition result (S10), as shown in example C in Fig. 9. Note that in step S10, the emotion selection unit 27 may assign one emotion label only to the overlapping portion of the speech section for voice recognition and the speech emotion recognition utterance section, but is not limited to this.
[0082] After step S9 or S10, the output unit 26 outputs the speech converted into text and the emotion label (S11), and the process ends.
[0083] The voice analysis system 1A according to the second embodiment described above can recognize multiple emotions and output multiple emotion labels even when multiple emotions are included in a single utterance. This allows the operator to analyze the multiple emotion labels that are output and determine what actions led to that emotion.
[0084] Furthermore, by integrating the results of voice recognition and voice emotion recognition and providing feedback to change the voice emotion recognition threshold for voice emotion recognition speech activity detection, voice emotion recognition can be performed with higher accuracy.
[0085] [Modification of the second embodiment] The voice recognition unit 23 of the voice analysis execution unit 20B may be replaced with an external voice recognition system. FIG. 11 is a block diagram showing an example of the configuration of a voice analysis execution unit 20C having an external voice recognition system.
[0086] The voice analysis execution unit 20C does not include the voice recognition unit 23 shown in Fig. 3. However, the voice analysis execution unit 20C causes an external voice recognition system 30 to perform voice recognition and voice-to-text processing. For this reason, the voice input unit 21 outputs voice data to the voice recognition system 30. The functions of the voice analysis execution unit 20C other than voice recognition are the same as those of the voice analysis execution unit 20B shown in Fig. 8.
[0087] The speech recognition system 30 performs speech recognition based on the speech data input from the speech input unit 21 of the speech analysis execution unit 20C, and converts the speech into text. The speech recognition system 30 then outputs the textual speech data to the output unit 26 of the speech analysis execution unit 20C. The textual speech data is accompanied by text and a timestamp.
[0088] The output unit 26 assigns an emotion label of the voice emotion recognized by the voice emotion recognition unit 25 to the text-converted voice data input from the voice recognition system 30 and outputs the data. The output unit 26 can assign the emotion label to the text-converted voice at the correct timing using a timestamp.
[0089] [Example of information display] Next, two display examples of information output by the output unit 26 according to the first and second embodiments will be described with reference to FIGS. FIG. 12 is an interface diagram showing a first display example of the voice analysis result.
[0090] The speech analysis result 100 shown in Fig. 12 is displayed by the output unit 26. The speech analysis result 100 is the result of detecting and recognizing the content of the conversation in real time. The output unit 26 outputs the speech analysis result 100 in which the text-converted speech is arranged in chronological order and an emotion label is written for each predetermined portion of the speech.
[0091] The speech analysis result 100 displays the converted speech in chronological order as a speech recognition result 101. For example, the speech recognition result 101 displays a conversation between speakers as a single segment. Note that when a conversation between the same speaker continues, the output unit 26 can appropriately change the predetermined segment, for example, to a conversation every 10 seconds or every two sentences.
[0092] For each speech recognition result 101, an emotion label assigned to the speech emotion recognition result 102 is displayed. Information such as a hold time 103 is also displayed between multiple speech recognition results 101. Any other information may be displayed in the speech analysis result 100.
[0093] FIG. 13 is an interface diagram showing a second example of displaying the voice analysis results.
[0094] The speech analysis result 110 shown in FIG. 13 may be displayed by the output unit 26. The speech analysis result 110 is a result of analyzing a dialogue in real time and recognizing changes in emotions over time. The output unit 26 outputs the speech analysis result 110 by arranging the converted text speech in chronological order and graphing emotion labels for each predetermined time. The output unit 26 may set the predetermined time as, for example, the time required for one conversation by the speaker, or the time required for two or three conversations.
[0095] The speech analysis result 110 displays speech recognition text 111, which correlates the dialogue time with the dialogue content, and speech emotion 112, which visualizes emotions that change over time. When multiple emotions are recognized simultaneously, speech emotion 112 indicates the strength of those emotions. For example, the closer the value is to "1.0," the stronger the emotion.
[0096] The voice analysis results may be output in the form of a bar, as shown in output (3) of Figure 9. In this case, the textual voice is not displayed, but it is possible to visually identify the emotions of the other person by, for example, assigning different colors to each emotion. The voice analysis results 100, 110 may also be the results of batch processing.
[0097] The present invention is not limited to the above-described embodiments and includes various modifications. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, or to add the configuration of another embodiment to the configuration of one embodiment. Furthermore, it is possible to add, delete, or replace part of the configuration of each embodiment with other configurations.
[0098] Furthermore, the above-described configurations, functions, processing units, etc. may be partially or entirely implemented in hardware, for example, by designing them as integrated circuits. The above-described configurations, functions, etc. may also be implemented in software, with a processor interpreting and executing a program that implements each function. Information such as the programs, tables, and files that implement each function can be stored in a memory, a recording device such as a hard disk or SSD (Solid State Drive), or a recording medium such as an IC card or DVD. In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0099] 1, 1A... Speech analysis system, 14... Speech input device, 15... Output device, 20, 20A, 20B, 20C... Speech analysis execution unit, 21... Speech input unit, 22... Speech recognition speech period detection unit, 23... Speech recognition unit, 24... Speech emotion recognition speech period detection unit, 25... Speech emotion recognition unit, 26... Output unit, 27... Emotion selection unit, 100, 110... Speech analysis result
Claims
1. an audio input unit for inputting an audio signal; a speech-period detection unit for speech recognition that detects a first speech period for speech recognition from the speech signal exceeding a speech-period threshold for detecting a first speech period for speech recognition, and extracts a first speech segment; a speech section detection unit for voice emotion recognition that detects a second speech section for voice emotion recognition from the speech signal that exceeds a threshold for voice emotion recognition that is higher than the speech section threshold, and extracts a second speech segment; a voice emotion recognition unit that performs voice emotion recognition processing based on the second utterance segment extracted in the second utterance period and assigns emotion to the second utterance segment; an output unit that associates the emotion with the voice that has been converted into text and outputs the voice that has been converted into text and the emotion. Voice analysis system.
2. The speech segment detection unit for voice emotion recognition sets the threshold for voice emotion recognition, divides the second speech segment into utterance units, and removes noise mixed in the voice. The speech analysis system of claim 1 .
3. an emotion selection unit that selects the emotion to be output by the output unit based on the emotion assigned to the second utterance section; The speech analysis system of claim 2 .
4. a speech recognition unit that performs speech recognition processing based on the first speech segment extracted in the first speech section, converts the speech into text, and outputs the converted text to the output unit; The speech analysis system of claim 3 .
5. The emotion selection unit assigns one emotion label to the second utterance section when the first utterance section and the second utterance section are the same. The speech analysis system of claim 4 .
6. When there is no result of the voice emotion recognition processing for the first utterance section, the emotion selection unit changes the voice emotion recognition threshold to a value higher than the previous value, and causes the voice recognition speech section detection unit to redetect the first utterance section. The speech analysis system of claim 5 .
7. When a result of the voice emotion recognition processing for the first utterance section exists and the confidence level of the result of the voice emotion recognition processing is lower than a confidence level threshold, the emotion selection unit changes the voice emotion recognition threshold to a value lower than the previous value and causes the voice emotion recognition utterance section detection unit to redetect the second utterance section. The speech analysis system of claim 6 .
8. When a plurality of results of the voice emotion recognition process exist for the first utterance section, the emotion selection unit integrates the plurality of emotion labels and assigns them to the first utterance section. The speech analysis system of claim 6 .
9. The output unit arranges the converted speech in chronological order and outputs a speech analysis result along with the emotion label for each predetermined portion of the speech. The speech analysis system of claim 7 .
10. The output unit arranges the converted speech in time series and outputs a speech analysis result in which the emotion labels are graphed at predetermined time intervals. The speech analysis system of claim 7 .
11. inputting an audio signal; detecting a first speech period for speech recognition from the speech signal exceeding a speech period threshold for detecting a first speech period for speech recognition, and extracting a first speech segment; detecting a second speech period for voice emotion recognition from the speech signal that exceeds a threshold for voice emotion recognition that is higher than the speech period threshold, and extracting a second speech segment; performing a voice emotion recognition process based on the second speech segment extracted in the second speech period, and assigning an emotion to the second speech segment; and associating the emotion with the voice that has been converted into text, and outputting the voice that has been converted into text and the emotion. Voice analysis methods.
Citation Information
Patent Citations
System and method for voice feeling recognition
JP1999119791A