Methods of estimating feelings

The emotion estimation method enhances emotion estimation by separating speech data into multiple vectors, using a learning model trained with specific loss functions, to accurately estimate expressed and intrinsic emotions, improving customer response and support quality.

JP2026056444APending Publication Date: 2026-04-01TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-19
Publication Date
2026-04-01

AI Technical Summary

Technical Problem

Existing emotion estimation technologies, such as those disclosed in Patent Document 1, do not effectively utilize machine learning to estimate customer emotions during business negotiations, limiting improvements in customer response and support quality.

Method used

An emotion estimation method that separates speech data into first and second vector data, using a learning model trained with a first loss function based on linguistic information, a second loss function for symmetric or asymmetric learning, and a third loss function to minimize mutual information, enabling the estimation of expressed and intrinsic emotions.

Benefits of technology

The method improves emotion estimation accuracy by estimating multiple emotions with different characteristics, including the emotion gap, achieving higher F1 scores for both expressed and intrinsic emotions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026056444000001_ABST
    Figure 2026056444000001_ABST
Patent Text Reader

Abstract

Improve emotion estimation technology. [Solution] An emotion estimation method performed by an information processing device, comprising: acquiring audio data; inputting the audio data into a learning model and separating it into at least first vector data and second vector data; and estimating an emotion corresponding to the audio data based at least on the first vector data and the second vector data, wherein the learning model is trained based on a first loss function based on the difference between linguistic information based on the audio data and the first vector data; a second loss function based on symmetric or asymmetric learning of the second vector data; and a third loss function that minimizes the mutual information between the first vector data and the second vector data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an emotion estimation method.

Background Art

[0002] Conventionally, techniques for analyzing the content of business negotiations are known. For example, Patent Document 1 discloses a dialogue analysis system that checks whether a salesperson explains matters to be explained and does not state matters that should not be stated in business negotiations with customers.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The success or failure of business negotiations may be related to the emotions of customers. Patent Document 1 does not disclose estimating the emotions of customers or the like by utilizing machine learning or the like. Also, not limited to business negotiations, if emotions can be estimated from voice data, it can lead to improvements in customer response, customer support quality, etc., but emotion estimation technology has not been sufficiently studied so far. Thus, there has been room for improvement in emotion estimation technology.

[0005] In view of such circumstances, an object of the present disclosure is to improve emotion estimation technology.

Means for Solving the Problems

[0006] An emotion estimation method according to one embodiment of the present disclosure is an emotion estimation method performed by an information processing device, comprising: acquiring speech data; inputting the speech data into a learning model and separating it into at least first vector data and second vector data; and estimating an emotion corresponding to the speech data based at least on the first vector data and the second vector data, wherein the learning model is trained on a first loss function based on the difference between linguistic information based on the speech data and the first vector data; a second loss function based on symmetric or asymmetric learning of the second vector data; and a third loss function that minimizes the mutual information between the first vector data and the second vector data. [Effects of the Invention]

[0007] According to one embodiment of this disclosure, emotion estimation technology is improved. [Brief explanation of the drawing]

[0008] [Figure 1] This is a block diagram showing the schematic configuration of the system according to this embodiment. [Figure 2] This is a flowchart showing the operation of an information processing device. [Modes for carrying out the invention]

[0009] The embodiments of this disclosure will now be described. Referring to Figure 1, the overview and configuration of System 1 according to this embodiment will be described. System 1 according to this embodiment comprises an information processing device 10 and a terminal device 20. The information processing device 10 is a server device installed, for example, in a data center. The terminal device 20 is any device used by each user. These devices are connected to each other so as to be able to communicate via a network 30 such as the Internet. In Figure 1, one information processing device 10 and one terminal device 20 are shown, but System 1 may comprise multiple such devices.

[0010] First, an overview of the emotion estimation technology according to this embodiment will be described, and details will be described later. The emotion estimation technology according to this embodiment is performed by an information processing device 10. Initially, the information processing device 10 acquires voice data such as business negotiations. The information processing device 10 inputs the voice data into a learning model and separates it into at least first vector data and second vector data. The information processing device 10 estimates the emotion corresponding to the voice data based at least on the first vector data and the second vector data. The learning model is trained based on a first loss function based on the difference between linguistic information based on the voice data and the first vector data, a second loss function based on symmetric or asymmetric learning of the second vector data, and a third loss function that minimizes the mutual information between the first vector data and the second vector data.

[0011] As described above, according to this embodiment, the information processing device 10 inputs speech data into a learning model and separates it into at least a first vector data and a second vector data. The learning model is trained based on a first loss function, a second loss function, and a third loss function. Therefore, according to this embodiment, emotion estimation is possible using at least two different vectors. Specifically, at least the first vector data and the second vector data can be used to estimate two different types of emotions: expressed emotion and intrinsic emotion. Expressed emotion is an emotion expressed through linguistic information. Intrinsic emotion is an emotion or sensation that resides in the mind and is not expressed as linguistic information. In other words, intrinsic emotion is an emotion expressed through at least one of paralinguistic information and nonverbal information. Linguistic information is information indicating the content of the utterance based on the speech data. Paralinguistic information is information such as emotion, attitude, and intention based on the speech data. Nonverbal information is information such as the speaker's age and gender based on the speech data. Thus, according to this embodiment, emotion estimation technology is improved in that it is possible to estimate multiple emotions with different characteristics, and furthermore, the differences between these emotions (hereinafter also referred to as the emotion gap) can also be estimated.

[0012] Next, the configurations of the information processing device 10 and the terminal device 20 will be described in detail. As shown in Figure 1, the information processing device 10 comprises a control unit 11, a storage unit 12, an input unit 13, an output unit 14, and a communication unit 15. The control unit 11 includes at least one processor. The processor is a general-purpose processor such as a CPU, or a dedicated processor specialized for a specific process. The control unit 11 controls each part of the information processing device 10 and executes processes related to the operation of the information processing device 10. The storage unit 12 includes at least one semiconductor memory, etc. The semiconductor memory is, for example, RAM or ROM. The storage unit 12 functions, for example, as a main memory or auxiliary memory. The storage unit 12 stores data used for the operation of the information processing device 10 and data obtained by the operation of the information processing device 10. For example, the storage unit 12 stores a learning model. The learning model is a model created by machine learning using a machine learning algorithm. The learning model may be a machine learning model built on a decision tree, a Convolutional Neural Network (CNN), a Recurrent Neural Network (RNN), or a model generated based on other deep learning algorithms. The input unit 13 includes at least one input interface. The input interface may be, for example, a physical key, a touchscreen, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts operations to input data used for the operation of the information processing device 10. The output unit 14 includes at least one output interface. The output interface may be, for example, a display that outputs information as video, or a speaker that outputs information as audio. The output unit 14 outputs data obtained by the operation of the information processing device 10. The communication unit 15 includes at least one external communication interface. The communication interface may be either a wired communication interface or a wireless communication interface. In the case of wired communication, the communication interface may be, for example, LAN or USB.In the case of wireless communication, the communication interface is, for example, an interface compatible with mobile communication standards such as 5G, or an interface compatible with short-range wireless communication. The communication unit 15 receives data used for the operation of the information processing device 10 and transmits data obtained by the operation of the information processing device 10.

[0013] As shown in Figure 1, the terminal device 20 comprises a control unit 21, a storage unit 22, an input unit 23, an output unit 24, and a communication unit 25. The control unit 21 includes at least one processor. The processor is a general-purpose processor such as a CPU, or a dedicated processor specialized for specific processing. The control unit 21 controls each part of the terminal device 20 and executes processing related to the operation of the terminal device 20. The storage unit 22 includes at least one semiconductor memory, etc. The semiconductor memory is, for example, RAM or ROM. The storage unit 22 functions, for example, as a main memory or auxiliary memory. The storage unit 22 stores data used for the operation of the terminal device 20 and data obtained by the operation of the terminal device 20. The input unit 23 includes at least one input interface. The input interface may be, for example, a physical key, a touchscreen, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 23 accepts operations to input data used for the operation of the terminal device 20. The output unit 24 includes at least one output interface. The output interface is, for example, a display that outputs information as video, or a speaker that outputs information as audio. The output unit 24 outputs data obtained by the operation of the terminal device 20. The communication unit 25 includes at least one external communication interface. The communication interface may be either a wired communication interface or a wireless communication interface. In the case of wired communication, the communication interface is, for example, a LAN or USB. In the case of wireless communication, the communication interface is, for example, an interface compatible with a mobile communication standard such as 5G, or an interface compatible with short-range wireless communication. The communication unit 25 receives data used in the operation of the terminal device 20 and transmits data obtained by the operation of the terminal device 20.

[0014] The functions of the information processing device 10 or terminal device 20 are realized by executing a program according to this embodiment on a processor corresponding to the control unit 11 or control unit 21. In other words, the functions of the information processing device 10 or terminal device 20 are realized by software. The program causes the computer to function as the information processing device 10 or terminal device 20 by having the computer execute the operations of the information processing device 10 or terminal device 20. In other words, the computer functions as the information processing device 10 or terminal device 20 by executing the operations of the information processing device 10 or terminal device 20 according to the program. In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-temporary computer-readable media, such as magnetic recording devices and semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending a portable recording medium such as a DVD on which the program is recorded. Alternatively, the program may be distributed by storing the program on the storage of an external server and transmitting the program from the external server to another computer. The program may also be provided as a program product. Some or all of the functions of the information processing device 10 or the terminal device 20 may be implemented by a dedicated circuit corresponding to the control unit 11 or the control unit 21. In other words, some or all of the functions of the information processing device 10 or the terminal device 20 may be implemented by hardware.

[0015] Referring to Figure 2, the operation of the information processing device 10 according to this embodiment will be described. First, the control unit 11 of the information processing device 10 acquires voice data (step S10). Any method can be used to acquire the voice data. For example, the control unit 11 may acquire voice data via the input unit 13. Alternatively, the control unit 11 may acquire voice data from external devices, including the terminal device 20, via the communication unit 15 and the network 30. The voice data includes the voice of a specific speaker, such as in a business negotiation. The voice data is not limited to this and may include any data, such as telephone conversations with customers. The specific speaker may be, for example, a customer or staff member in a business negotiation. The business negotiation may be, for example, a negotiation related to the sale of a vehicle.

[0016] Next, the control unit 11 inputs the audio data into the learning model and separates it into at least first vector data and second vector data (step S20). The learning model is trained based on a first loss function based on the difference between the linguistic information based on the audio data and the first vector data, a second loss function based on symmetric or asymmetric learning of the second vector data, and a third loss function that minimizes the mutual information between the first and second vector data. In other words, the learning model is trained by the first loss function to minimize the difference between the linguistic information based on the audio data and the first vector data. The learning model is also trained by a second loss function based on symmetric or asymmetric learning of the second vector data, depending on the symmetry or asymmetry between the data. If the second loss function is a function based on symmetric learning, the symmetric learning may be simCLR. In this case, identical audio is treated as a positive sample, and different audio is treated as a negative sample. If the second loss function is a function based on asymmetric learning, the asymmetric learning may be BYOL, SimSiam, or DINO. In this case, the student input is an audio segment shorter than a predetermined length, and the teacher input is an audio segment longer than a predetermined length. The learning model is trained based on a third loss function that minimizes the mutual information between the first and second vector data and separates them as far apart as possible. CLUB or DiCy may be used for training related to the third loss function. The language information based on the audio data may be transcript data related to the audio data. Any method may be used to generate the transcript data.

[0017] Next, the control unit 11 estimates the emotion corresponding to the voice data based at least on the first vector data and the second vector data (step S30). For example, the control unit 11 may input the first vector data and the second vector data into a first estimation model and a second estimation model, respectively, to estimate the emotion. The first estimation model and the second estimation model may be, for example, a logistic regression or an ECAPA-TDNN model.

[0018] Subsequently, the control unit 11 outputs an estimation result of the emotion corresponding to the voice data (step S40). Any method can be adopted for the output process of the estimation result. For example, the control unit 11 may transmit data related to the estimation result to the terminal device 20 via the communication unit 15, and the output unit 24 of the terminal device 20 may output the estimation result. The control unit 21 may output the estimation result through a user interface displayed and output by the output unit 24.

[0019] Note that, for a part of the teacher data used when training the first estimation model and the second estimation model, pseudo labels may be attached instead of labels, and the first estimation model and the second estimation model may be trained by semi-supervised learning. For example, pseudo labels may be attached to more than half of the teacher data instead of labels. Also, the first estimation model may be trained by supervised learning, and only the second estimation model may be trained by semi-supervised learning. Here, by training the learning model with the first loss function, the first vector data becomes data corresponding to the language information of the voice data. On the other hand, by the second loss function and the third loss function, the second vector data is adjusted so that the correlation with the first vector data becomes low. In other words, the second vector data becomes data corresponding to information other than the language information of the voice data (paralinguistic information and non-verbal information). Here, the diversity related to paralinguistic information and non-verbal information is low, and it is considered that learning can be performed with a relatively small number of data.

[0020] As described above, the information processing device 10 inputs voice data into the learning model, separates it into at least the first vector data and the second vector data, and estimates the emotion corresponding to the voice data based at least on the first vector data and the second vector data.

[0021] According to such a configuration, the information processing device 10 can estimate emotions using at least two different vectors. In this way, the emotion estimation technology is improved in that emotions with different properties can be estimated and the emotion gap can be estimated.

[0022] Table 1 shows the accuracy verification results of the first estimation model and the second estimation model. Here, for the training data, labels of expressed emotions (hereinafter also referred to as language labels) and labels of inherent emotions (hereinafter also referred to as psychological labels) are attached, and the accuracy verification of the first estimation model and the second estimation model is carried out using the first vector data and the second vector data as inputs respectively. When the language label and the psychological label do not match, the F1 score of the psychological label is improved by 6 points when using the second vector data compared to using the first vector data. On the other hand, the F1 score of the language label is improved by 7 points when using the first vector data compared to using the second vector data. For example, when the language label and the psychological label are different, by using the estimation result with a higher F1 score, an emotion estimation result with high accuracy can be provided.

[0023]

Table 1

[0024] Although the present disclosure has been described based on the drawings and embodiments, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions etc. included in each component or each step etc. can be rearranged so as not to be logically contradictory, and it is possible to combine or divide a plurality of components or steps etc. into one.

[0025] For example, in the above-described embodiment, an embodiment in which the configurations and operations of the information processing apparatus 10 or the terminal apparatus 20 are distributed to a plurality of computers capable of communicating with each other is also possible.

[0026] Furthermore, in this embodiment, for example, audio data is input to a learning model, separated into at least two vectors, and two emotions with different properties, expressed emotion and intrinsic emotion, are estimated using the first and second vector data, respectively. However, the embodiment is not limited to this. For example, audio data may be input to a learning model and separated into three vector data. In this case, for example, the three vector data may be used to estimate emotion based on linguistic information, emotion based on paralinguistic information, and emotion based on nonverbal information, respectively. [Explanation of Symbols]

[0027] 10: Information processing device, 11: Control unit, 12: Storage unit, 13: Input unit, 14: Output unit, 15: Communication unit, 20: Terminal device, 21: Control unit, 22: Storage unit, 23: Input unit, 24: Output unit, 25: Communication unit, 30: Network

Claims

1. A method for estimating emotions performed by an information processing device, Acquiring audio data and, The aforementioned audio data is input into a learning model and separated into at least a first vector data and a second vector data, Estimating the emotion corresponding to the voice data based at least on the first vector data and the second vector data, A method for estimating emotions, including, A method for estimating emotion, wherein the learning model is trained based on a first loss function based on the difference between language information based on the speech data and the first vector data, a second loss function based on symmetric or asymmetric learning of the second vector data, and a third loss function that minimizes the mutual information between the first vector data and the second vector data.

2. A sentiment estimation method according to claim 1, wherein the second loss function is a function based on symmetric learning, and the symmetric learning includes simCLR.

3. A sentiment estimation method according to claim 1, wherein the second loss function is a function based on asymmetric learning, and the asymmetric learning includes BYOL, SimSiam, or DINO.

4. A method for estimating emotions according to claim 1, A sentiment estimation method in which CLUB or DiCy is used in training related to the third loss function.

5. A method for estimating emotions according to claim 1, The aforementioned language information is a transcription data relating to the aforementioned audio data, in an emotion estimation method.

Citation Information

Patent Citations

  • Dialogue analysis system and dialogue analysis program

    JP2019028910A