Emotion inference system

The emotion estimation system addresses the challenge of inaccurate emotion recognition by enhancing voice features through data compression, restoration, and low-sampling rates, achieving high accuracy and intuitive visual output.

JP2026014637APending Publication Date: 2026-01-29SAITAMA UNIVERSITY +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115984
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional emotion recognition methods struggle to accurately identify emotions with high probability, especially in short utterances or poor acoustic environments, due to limitations in enhancing voice characteristics.

Method used

An emotion estimation system that enhances voice features through machine learning using voice data compression and restoration, feature extraction, and low-sampling rates, along with speaker separation for multiple speakers, and outputs emotions visually using light colors.

Benefits of technology

The system achieves high-probability emotion estimation by enhancing voice features, improving accuracy and communication speed, and allows intuitive visual recognition of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014637000001_ABST
    Figure 2026014637000001_ABST
Patent Text Reader

Abstract

To provide an emotion estimation system capable of estimating an emotion with high probability by enhancing features of a voice itself.SOLUTION: An emotion estimation system is a system for estimating an emotion of an object speaker from a voice uttered from the object speaker to be an estimation object, and includes a voice analysis model 1, a voice data acquisition part 2, and an emotion output part 3. The voice analysis model 1 performs the machine learning process using feature voice data obtained by performing a feature extraction process of dividing a frequency band of voice data such that the number of divisions gradually decreases to compress the voice data and restoring the compressed voice data such that the number of divisions gradually increases to extract a feature of the voice data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an emotion estimation system that estimates the emotion of a target speaker from speech uttered by the target speaker. [Background technology]

[0002] Conventionally, as shown in Patent Document 1 below, an emotion recognition method and apparatus for this type of emotion estimation includes a step of extracting a set of at least one feature derived from a speech signal and a step of processing the extracted set of features to detect the emotion conveyed by the speech signal, and further includes a step of processing the speech signal with a low-pass filter before extracting the at least one feature of the set from the speech signal. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-99084 Summary of the Invention [Problem to be solved by the invention]

[0004] Specifically, such conventional emotion recognition methods and devices aim to accurately identify emotions with a light workload, even for short utterances or utterances made in poor acoustic environments, by filtering the voice signal strength at a cutoff frequency (Fco) in the low-pass filter processing step, for example, in a range of 150 to 400 Hz. However, there are limitations to filtering, and the characteristics of the voice itself are not enhanced to estimate emotions with a high probability.

[0005] The present invention has been made in view of the above background, and has an object to provide an emotion estimation system that can estimate emotions with a high probability by enhancing the features of the voice itself. [Means for solving the problem]

[0006] The emotion estimation system of the first invention comprises: An emotion estimation system that estimates the emotion of a target speaker from a speech uttered by the target speaker, the emotion estimation system comprising: A voice analysis model constructed by machine learning processing using speech data of a voice uttered by a speaker and the correct emotion of the voice as training data; a voice data acquisition unit that acquires voice data of the voice of the target speaker; an emotion output unit that outputs an emotion estimated by inputting the voice data acquired by the voice data acquisition unit into the voice analysis model; Equipped with The voice analysis model is characterized in that the voice data is compressed by dividing the frequency band so that the number of divisions gradually decreases, and the machine learning processing is performed using characteristic voice data that has undergone feature extraction processing, in which the features of the voice data are extracted by restoring the compressed voice data so that the number of divisions gradually increases.

[0007] According to the emotion estimation system of the first invention, when constructing a voice analysis model through machine learning processing using voice data of a voice uttered by a speaker and the correct emotion of the voice as training data, the voice data is divided into frequency bands and compressed so that the number of divisions gradually decreases, and the voice data compressed so that the number of divisions gradually increases is restored to extract feature voice data that has been subjected to feature extraction processing to extract features of the voice data. Unlike noise removal by filter processing, this makes it possible to extract features necessary for emotion estimation.

[0008] In this way, the emotion estimation system of the first aspect of the present invention can estimate emotions with a high probability by enhancing the features of the voice itself.

[0009] The emotion estimation system of the second invention is the emotion estimation system of the first invention, The voice analysis model is characterized in that when performing the machine learning processing using the feature voice data that has been repeatedly subjected to the feature extraction processing, the machine learning processing is performed using fully connected voice data that has been fully connected after excluding a portion of the compressed voice data or the voice data to be restored.

[0010] According to the emotion estimation system of the second invention, when machine learning processing is performed using feature audio data obtained by repeatedly performing feature extraction processing on audio data, which involves compressing and restoring the audio data by dividing the frequency bands, part of the compressed audio data or the audio data to be restored is excluded, and then the machine learning processing is performed using fully connected audio data. This makes it possible to intentionally remove part of the data and then perform learning again, thereby further increasing the probability of emotion estimation.

[0011] In this way, according to the emotion estimation system of the second invention, the features of the voice itself can be enhanced, and emotions can be estimated with a higher probability.

[0012] The emotion estimation system of the third invention is the emotion estimation system of the first invention, The emotion output unit is characterized in that it outputs an emotion estimated by inputting low-sampled audio data obtained by sampling the audio data acquired by the audio data acquisition unit at a low sampling rate of 3.0 kHz or more and less than 10.0 kHz into the audio analysis model.

[0013] According to the emotion estimation system of the third invention, by inputting low-sampled speech data obtained by sampling speech data of a target speaker to be estimated at a low sampling rate of less than 10.0 kHz into a speech analysis model configured by machine learning using feature speech data that has been subjected to feature extraction processing, the emotion estimation system can estimate and output emotions with high accuracy even with low-sampled speech data, which has a small amount of data and contributes to improving communication speed and analysis speed.

[0014] As described above, according to the emotion estimation system of the third invention, the speech data of the target speaker is converted into low-sampled data with a small amount of data, while the features of the speech itself are enhanced, making it possible to estimate emotions with a high probability.

[0015] The emotion estimation system of the fourth invention is the emotion estimation system of the first invention, When there are multiple target speakers to be estimated, the voice data acquisition unit acquires voice data of the voices uttered by the multiple target speakers, and the emotion output unit separates the acquired voice data into voice data for each target speaker, and inputs the separated voice data into the voice analysis model to output an estimated emotion for each target speaker.

[0016] According to the emotion estimation system of the fourth invention, when multi-person speech is generated from multiple target speakers to be estimated, the speech is classified by first person and then input into a speech analysis model, thereby making it possible to simultaneously estimate the emotions of each target speaker.

[0017] As described above, according to the emotion estimation system of the fourth aspect of the present invention, even when there are multiple target speakers, the emotion of each target speaker can be estimated with a higher probability by improving the features of the speech itself.

[0018] A feeling estimation system according to a fifth aspect of the present invention is any one of the first to fourth aspects of the present invention, The emotion output unit outputs the estimated emotion as a light color corresponding to the emotion.

[0019] According to the emotion estimation system of the fifth aspect of the present invention, the emotion estimation result is presented by, for example, changing the light color depending on the type of emotion. This makes it possible to output the estimation result using light color, which is an intuitive method of recognizing the result as visual information.

[0020] As described above, the emotion estimation system of the fifth aspect of the present invention can estimate emotions with a high probability by enhancing the features of the voice itself, and output the estimation results in a way that is easy to intuitively recognize. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is a system configuration diagram showing the overall configuration of a feeling estimation system according to an embodiment of the present invention. [Figure 2] FIG. 2 is an explanatory diagram showing the processing content of the voice analysis model of FIG. 1. [Figure 3A] FIG. 2 is an explanatory diagram showing the processing content of the voice analysis model of FIG. 1. [Figure 3B] FIG. 2 is an explanatory diagram showing the processing content of the voice analysis model of FIG. 1. [Figure 4] FIG. 2 is an explanatory diagram showing the result of emotion estimation accuracy in the emotion estimation system of FIG. 1. [Figure 5] FIG. 2 is a system configuration diagram showing a modified example of the emotion estimation system of FIG. 1. DETAILED DESCRIPTION OF THE INVENTION

[0022] An emotion estimation system according to an embodiment of the present invention will be described below with reference to FIG.

[0023] As shown in FIG. 1, the emotion estimation system is a system that estimates the emotion of a target speaker from the speech uttered by the target speaker, and includes a speech analysis model 1, a speech data acquisition unit 2, and an emotion output unit 3.

[0024] The voice analysis model 1 is constructed by machine learning processing using multiple voice data of voices uttered by speakers and the correct emotions for each voice as training data. The voice data of voices uttered by speakers and the correct emotions for each voice as training data may be stored in advance in a voice database or may be obtained as updated via an external server or the like.

[0025] The voice data acquisition unit 2 acquires voice data of the voice of the target speaker to be estimated, and may be configured, for example, as in this embodiment, with a microphone 2 (schematically shown at the bottom of Figure 1), and may acquire not only the voice output of the voice of the target speaker output from the microphone 2 converted into voice data (voice file data), but also recorded voice data (voice file data) from an external server, recording medium, etc.

[0026] The emotion output unit 3 is a means for outputting an emotion estimated by inputting the voice data (voice file data) acquired by the voice data acquisition unit 2 into the voice analysis model 1, and the output means may be, for example, a light-emitting means such as a light-emitting diode 3 (schematically shown at the bottom of FIG. 1) that outputs an emotion as colored light as in this embodiment, or an audio output means such as a speaker that visually outputs the estimated emotion as characters or figures on a display means such as a panel, or that reads out the estimated emotion.

[0027] Next, with reference to FIG. 2, details of the voice analysis model 1, which is a feature of the emotion estimation system, will be described.

[0028] The voice analysis model 1 performs feature extraction processing on multiple voice data acquired by the voice data acquisition unit 2 (each of the multiple voice data is assigned an emotion label of the correct emotion whose corresponding emotion is known), and then performs machine learning processing using this as training data.

[0029] In this embodiment, the machine learning process, particularly deep learning, is performed by training a neural network for emotion estimation and comparing the outputs of multiple channels of the neural network with correct emotions (emotion labels), but this is not limited to this, and machine learning other than deep learning may also be adopted.

[0030] Here, the feature extraction process involves compressing the audio data by dividing the frequency band so that the number of divisions gradually decreases (such as in the figure, 128 → 112 → 80 → 64), and then restoring the compressed audio data so that the number of divisions gradually increases (such as in the figure, 64 → 80 → 128), thereby extracting the features of the audio data.

[0031] In this way, the audio data is compressed by dividing the frequency band so that the number of divisions becomes smaller in stages, and the compressed audio data is restored so that the number of divisions becomes larger in stages, and by using feature audio data that has been subjected to feature extraction processing to extract the features of the audio data, it is possible to extract the features necessary for emotion estimation, unlike noise removal using filter processing.

[0032] Here, it is preferable to repeat the feature extraction process multiple times before performing machine learning processing. In such multiple feature extraction processes, the machine learning processing is performed using fully connected audio data after excluding a portion of the compressed audio data or the audio data to be restored (Dropout in the figure).

[0033] The feature extraction process compresses the audio data by dividing the frequency band so that the number of divisions becomes smaller in stages (for example, in the figure, divisions 128 → 112 → 80 → 64). There are no particular restrictions on the stages of data compression and they can be selected as appropriate, but it is preferable to have at least two stages, and more preferably three stages, of data compression from the input audio data in order to extract features while removing information that may cause unnecessary noise.

[0034] By adding more convolutional and pooling layers (1D Conv+Pooling in the diagram), the model can extract higher-level, abstract features from the input data. This makes it possible to capture complex patterns and relationships, and is expected to improve estimation accuracy. However, increasing the number of convolutional and pooling layers increases training and inference times, requires more computing resources, and increases the complexity of the speech analysis model, increasing the risk of overfitting.

[0035] In this way, by performing machine learning processing using characteristic audio data that has been repeatedly subjected to feature extraction processing in which audio data is compressed and restored by dividing frequency bands, the accuracy of emotion estimation can be improved.

[0036] Furthermore, by excluding a portion of the compressed audio data or the audio data to be restored and then performing machine learning processing using the fully connected audio data, it is possible to intentionally remove some of the data and then retrain it, thereby avoiding the risk of overfitting in machine learning and further increasing the reliability of the results obtained by emotion estimation.

[0037] In fact, the emotion estimation accuracy (the rate of agreement between the correct emotion and the estimated result) of speech analysis model 1, which uses machine learning (deep learning) and employs the feature extraction processing structure shown in Figure 2, achieved a high emotion estimation probability of over 99%.

[0038] Here, the inventors of the present application verified that differences in the structure of the feature extraction process significantly affect the results of deep learning, as shown in Figure 3. Specifically, in order to search for a structure with high emotion estimation accuracy, they devised multiple feature extraction process structures and conducted experiments in which they repeated deep learning.

[0039] The emotion estimation accuracy was 62.8% for the feature extraction processing structure in Figure 3A, and 94.7% for the feature extraction processing structure in Figure 3B. The learning conditions were the same: learning rate 0.01, loss function categorical crossentropy, optimization method Adam, number of trials 30, and batch size 2048.

[0040] Next, the processing in emotion output section 3 will be described in detail.

[0041] The emotion output unit 3 converts the voice data acquired by the voice data acquisition unit 2 into low-sampled voice data sampled at a low sampling rate, and then inputs the low-sampled voice data into the voice analysis model 1 to output an estimated emotion.

[0042] More specifically, low-sampled audio data obtained by sampling the audio data acquired by the audio data acquisition unit 2 at a low sampling rate of 3.0 kHz or more and less than 10.0 kHz is used.

[0043] The inventors of the present application therefore verified the emotion estimation accuracy at a low sampling rate of 6 kHz, as shown in Figure 4. Specifically, they compared the emotion estimation accuracy at 6 kHz with that at 16 kHz and 11 kHz, which are currently standard frequencies used in character recognition and emotion estimation. As a result, as shown in Figure 4, it was found that the emotion estimation accuracy at 6 kHz was comparable to that at 16 kHz and 11 kHz. Note that emotion estimation accuracy was measured by using a binomial logistic regression model constructed from training data to estimate emotions on test data containing correct emotion information, and then comparing the estimated results with the correct emotions to calculate the match rate.

[0044] Furthermore, in this embodiment, the emotion output unit 3 uses a light-emitting diode 3 that outputs emotion-indicating colored light to output the emotion. Therefore, the emotion estimation result is presented with a light color that changes depending on the type of emotion. This makes it possible to output the estimation result using a light color that is intuitively easy to recognize as visual information.

[0045] Here, there is no particular limitation on the light color presented according to the type of emotion, but presenting a color or light color that people associate with emotions, such as red representing anger, makes it easier for the emotion output unit 3 to intuitively recognize the output of the estimated emotion as visual information.

[0046] The color of light to be output may be appropriately selected, for example, to be a color corresponding to an emotion listed in color psychology or the Plutchik Wheel of Emotions.

[0047] The light emitting diode 3 that outputs emotions as colored light may be a light source unit equipped with multiple types of light emitting diodes that emit different monochromatic light, or may be a light source unit equipped with a full-color LED (3-in-1 LED or RGB-LED), and the light color can be changed depending on the type of emotion to present the result of emotion estimation.

[0048] The above is the details of the emotion estimation system of the present embodiment. According to this emotion estimation system, by using feature audio data that has been subjected to feature extraction processing that extracts features of audio data, it is possible to extract features necessary for emotion estimation, unlike noise removal by filter processing, and to enhance the features of the audio itself, thereby enabling emotion estimation with a high probability.

[0049] Next, a modification of the above-described emotion estimation system will be described with reference to FIG.

[0050] Specifically, when there are multiple target speakers to be estimated, the speech data acquisition unit 2 acquires speech data of speech uttered by the multiple target speakers (speech in which the voices of the multiple target speakers are mixed).

[0051] Therefore, the emotion output unit 3 performs a speaker separation process to separate the acquired voice data for each target speaker, and then inputs the separated voice data into the voice analysis model 1 to output an estimated emotion for each target speaker.

[0052] The speaker separation process may employ various methods for speaker separation, such as using prepared speech data to train a neural network for speaker separation through machine learning (deep learning) and comparing the outputs of multiple channels of the neural network with the correct speech.

[0053] In this way, for multi-person speech with multiple target speakers to be estimated, by classifying it by first person and then inputting it into the speech analysis model, it is possible to simultaneously estimate the emotions of each target speaker. [Explanation of symbols]

[0054] 1...Voice analysis model 2...Voice data acquisition unit 3...Emotion output unit

Claims

1. An emotion estimation system that estimates the emotion of a target speaker from a speech uttered by the target speaker, the emotion estimation system comprising: A voice analysis model constructed by machine learning processing using speech data of a voice uttered by a speaker and the correct emotion of the voice as training data; a voice data acquisition unit that acquires voice data of the voice of the target speaker; an emotion output unit that outputs an emotion estimated by inputting the voice data acquired by the voice data acquisition unit into the voice analysis model; Equipped with The voice analysis model performs the machine learning process using feature voice data that has undergone feature extraction processing, in which the voice data is divided into frequency bands and compressed so that the number of divisions gradually decreases, and the compressed voice data is restored so that the number of divisions gradually increases, thereby extracting features of the voice data.

2. The emotion estimation system according to claim 1, The emotion estimation system is characterized in that, when performing the machine learning processing using the feature voice data obtained by repeatedly performing the feature extraction processing, the voice analysis model performs the machine learning processing using fully connected voice data that is fully connected after excluding a portion of the compressed voice data or the voice data to be restored.

3. The emotion estimation system according to claim 1, the emotion output unit outputs an estimated emotion by inputting low-sampled audio data obtained by sampling the audio data acquired by the audio data acquisition unit at a low sampling rate of 3.0 kHz or more and less than 10.0 kHz to the audio analysis model.

4. The emotion estimation system according to claim 1, an emotion estimation system, wherein, when there are multiple target speakers to be estimated, the voice data acquisition unit acquires voice data of voices uttered by the multiple target speakers, and the emotion output unit separates voice data for each target speaker from the acquired voice data, and inputs the separated voice data to the voice analysis model to output an emotion estimated for each target speaker.

5. The emotion estimation system according to any one of claims 1 to 4, The emotion estimation system, wherein the emotion output unit outputs the estimated emotion as a light color corresponding to the emotion.

Citation Information

Patent Citations

  • Emotion recognition method and device

    JP2003099084A