Emotion recognition system and emotion recognition method
The emotion recognition system addresses subjective deviations by inferring differential emotions from the same individual's voice data, enhancing accuracy by analyzing relative emotional changes and reducing errors from speaker and environmental factors.
Patent Information
- Application Number
- JP2021183862
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-11-11
AI Technical Summary
Existing emotion recognition systems rely on subjective judgments from a small number of developers and labelers, leading to deviations from actual user impressions, especially when considering speaker and environmental variations.
An emotion recognition system that infers differential emotions by comparing voice data from the same individual, using a differential emotion recognition model to analyze the relative emotional changes in two pieces of voice data, reducing the influence of speaker and environmental characteristics.
The system accurately recognizes emotions by focusing on relative emotional changes, aligning more closely with actual user impressions and reducing errors due to speaker and environmental variations.
Smart Images

Figure 0007706340000004 
Figure 0007706340000005 
Figure 0007706340000006
Abstract
Description
Technical Field
[0001] The present invention generally relates to a technique for inferring emotions expressed in speech.
Background Art
[0002] Emotion information contained in human speech plays an important role in human communication. Since it is possible to determine whether the purpose of communication has been achieved based on the movement of emotions, there is a need to analyze emotions in communication. In business scenes such as daily business activities and responses at call centers, analysis of emotions in many voice communications is required, so voice emotion recognition by machines is desired.
[0003] Voice emotion recognition outputs, for the input voice, the category of emotion contained in the voice, or the degree for each emotion category. As its mechanism, there are a method of classifying or performing regression analysis from the features of the voice signal based on predetermined rules, and a method of obtaining the rules by machine learning.
[0004] In recent years, a voice dialogue device capable of easily performing emotion recognition according to a user has been disclosed (see Patent Document 1).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] In the rule-based method, the result output based on one piece of voice data is determined by rules or thresholds defined by developers. Also, in the machine learning method, parameters are determined based on the voice data collected for learning and the label results by a person (labeler) who labels the emotional impressions received by listening to the voice data as correct values. However, in any case, since judgments are made depending on the subjectivity of a small number of people such as developers and labelers, the output of the emotion recognizer may deviate from the impressions of actual users.
[0007] The present invention has been made in consideration of the above points, and intends to propose an emotion recognition system and the like that can appropriately recognize the emotion expressed in voice.
Means for Solving the Problem
[0008] In order to solve such a problem, in the present invention, an input unit that inputs first voice data and second voice data, and a differential emotion recognition model that infers the differential emotion in the two pieces of voice data are input with the first voice data and the second voice data, and a processing unit that acquires differential emotion information indicating the differential emotion in the first voice data and the second voice data from the differential emotion recognition model are provided.
[0009] In the above configuration, for example, since the differential emotion (that is, the relative emotion) is inferred from two pieces of voice data, the emotion recognized by the emotion recognition system can be made closer to the impression of the actual user than the emotion inferred from one piece of voice data (that is, the absolute emotion).
Effect of the Invention
[0010] According to the present invention, an emotion recognition system that can accurately recognize the emotion expressed in voice can be realized. Problems, configurations, and effects other than the above will be clarified by the description of the following embodiments.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Best Mode for Carrying Out the Invention
[0012] (I) First Embodiment Hereinafter, an embodiment of the present invention will be described in detail. However, the present invention is not limited to the embodiment.
[0013] In the conventional technology, due to the influence of factors such as the variation factors of the speaker characteristics of the input voice and the variation factors of the environmental characteristics of the input voice, there is a possibility that an unintended result may be output because the rules do not function correctly. Recently, with the emergence of deep learning, in machine learning, it has become possible to handle more complex rules, and although efforts have been widely made to solve this problem and the accuracy has been improved, it cannot be said that the problem has been sufficiently solved.
[0014] In this regard, the emotion recognition system according to the present embodiment recognizes voice emotions by comparing voice data within an individual instead of comparing voice data between individuals. Voice emotion refers to emotions such as joy, anger, sorrow, and happiness, negative or positive, which are the expressions of a person's inner self through their voice. This emotion recognition system does not use the emotion labels assigned by listening to the voice of a single person (absolute evaluation labels in which emotions such as a person being happy or sad at present are digitized), but uses the emotion labels assigned by listening to two voices of the same person (difference evaluation (relative evaluation) labels in which emotions grasped from two voices of a person are digitized). Therefore, it is easier for the labeler to label, and the difference evaluation label is more reliable than the absolute evaluation label. Note that the two voices used when this emotion recognition system infers emotions are voices of the same speaker, but they do not necessarily have to be consecutive voices. However, it is preferable to use voices acquired on the same day and / or at the same location.
[0015] According to this emotion recognition system, it is possible to recognize voice emotions based on the difference in voices of the same person while suppressing the influence of speaker characteristics and environmental characteristics more than before.
[0016] Next, embodiments of the present invention will be described with reference to the drawings. The following description and drawings are examples for explaining the present invention, and for clarity of explanation, appropriate omissions and simplifications have been made. The present invention can be implemented in various other forms. Unless otherwise limited, each component may be in a single or plural number. In the following description, the same reference numerals are given to the same elements in the drawings, and the description thereof will be omitted as appropriate.
[0017] Note that notations such as "first", "second", "third", etc. in this specification and the like are attached to identify components, and do not necessarily limit numbers or orders. Also, the numbers for identifying components are used for each context, and the numbers used in one context do not necessarily indicate the same configuration in other contexts. Also, it does not prevent a component identified by a certain number from having the functions of a component identified by another number.
[0018] FIG. 1 is a diagram showing an example of a processing flow related to the differential emotion recognition device 101 of the present embodiment.
[0019] First, in the learning phase 110, the user 102 prepares learning voice data 111 and learning differential emotion label data 112. Next, the user 102 uses the differential emotion recognition device 101 to generate a differential emotion recognition model 113 by learning.
[0020] Next, in the inference phase 120, the user 102 inputs continuous voice data 121 into the differential emotion recognition device 101 and obtains emotion transition data 122.
[0021] FIG. 2 is a diagram showing an example (learning voice table 200) of the learning voice data 111.
[0022] The learning voice table 200 stores a plurality of voice waveforms (voice data). For each of the plurality of voice waveforms, a voice ID and a speaker ID are assigned. The voice ID is a code that uniquely identifies the voice waveform. The speaker ID is a code assigned to the speaker of the voice waveform and is a code that uniquely identifies the speaker. Additionally, the learning voice table 200 stores a plurality of voice waveforms of a plurality of persons.
[0023] FIG. 3 is a diagram showing an example (learning differential emotion label table 300) of the learning differential emotion label data 112.
[0024] The learning differential emotion label table 300 stores a plurality of differential emotions. The differential emotion is a label assigned by a labeler and is a label obtained by quantifying the emotion of the voice (second voice) of the voice waveform with the second voice ID based on the voice (first voice) of the voice waveform with the first voice ID. The first voice ID and the second voice ID indicate the voice waveforms corresponding to the voice IDs of the learning voice data 111.
[0025] The first voice ID and the second voice ID shall be those of the same speaker ID. The emotional labels (difference values) in the two inputs are labeled by the labeler. Note that, as in the prior art, the absolute value of the emotion for one voice is labeled (stored in the learning voice data 111), and during learning, the difference in the absolute values may be used as the difference value. For example, if the absolute value "0.1" indicating the emotion of the voice with voice ID "1" and the absolute value "0.2" indicating the emotion of the voice with voice ID "2" are stored in the learning voice table 200, the difference value "0.1" may be calculated as the differential emotion corresponding to the absolute value "0.1" of the first voice ID and the absolute value "0.2" of the second voice ID.
[0026] Also, the learning differential emotion label table 300 may store the differential emotions by a plurality of labelers. In that case, in learning, statistical values such as the average value of the differential emotions by a plurality of labelers are used. Also, the emotion category may be plural instead of one. In that case, the learning differential emotion label table 300 stores, as the differential emotion, a vector value instead of a scalar value.
[0027] FIG. 4 is a diagram showing an example of the configuration of the differential emotion recognition device 101.
[0028] The differential emotion recognition device 101 includes, as components, a storage device 400, a CPU 401, a display 402, a keyboard 403, and a mouse 404, similar to the configuration of a general PC (Personal Computer). Each component can transmit and receive data via a bus 405.
[0029] The storage device 400 includes, as programs, a learning program 411 and a differential emotion recognition program 421. These programs are read and executed by the CPU 401 by an OS (Operating System) (not shown) existing in the storage device 400 at startup.
[0030] The differential emotion recognition model 113 is, for example, a neural network in which the number of states in the input layer is a total of "1024" states, which is the sum of the feature amount "512" of the first voice and the feature amount "512" of the second voice, the hidden layer has one layer and the number of states is "512" states, and the number of states in the output layer is "1" state. For the input {x i (i = 1 ··· 1024)} to the input layer, the values {h j (j = 1 ··· 512)} of the hidden layer are calculated by (Equation 1).
Equation
[0031] The output y of the output layer is the differential value of the emotion of the second voice with respect to the first voice, and is calculated by (Equation 2).
Equation
[0032] Here, s is an activation function, for example, a sigmoid function, W is a weight, and b is a bias.
[0033] As a means for obtaining feature amounts from the first voice and the second voice, statistics etc. for time-series LLD (Low-Level Descriptors) described in the following Document 1 can be used. Document 1: Suzuki, "Recognition of Emotions Contained in Voice", Journal of the Acoustical Society of Japan, Vol. 71, No. 9 (2015), pp. 484-489
[0034] Note that this embodiment does not limit the structure of the neural network, and any neural network structure and activation function may be used. Also, the differential emotion recognition model 113 is not limited to a neural network, and any model may be used.
[0035] The functions of the differential emotion recognition device 101 (such as the learning program 411 and the differential emotion recognition program 421) may be realized, for example, by the CPU 401 reading a program from the storage device 400 and executing it (software), or by hardware such as a dedicated circuit, or by a combination of software and hardware. Note that one function of the differential emotion recognition device 101 may be divided into multiple functions, or multiple functions may be combined into one function. For example, the differential emotion recognition program 421 may include an input unit 422, a processing unit 423, and an output unit 424. Also, a part of the functions of the differential emotion recognition device 101 may be provided as another function, or may be included in other functions. Further, a part of the functions of the differential emotion recognition device 101 may be realized by another computer connectable to the differential emotion recognition device 101. For example, the learning program 411 may be provided in a first PC, and the differential emotion recognition program 421 may be provided in a second PC.
[0036] FIG. 5 is a diagram showing an example of a flowchart of the learning program 411.
[0037] First, the learning program 411 assigns initial values to the parameters W ij 1 , b i 1 , W i 2 , b i 2 of the differential emotion recognition model 113 (S501). As the initial values, the learning program 411 gives random values to facilitate the learning of the neural network.
[0038] Next, the learning program 411 reads data from the learning voice data 111 and the learning differential emotion label data 112 (S502).
[0039] Next, the learning program 411 updates the parameters of the differential emotion recognition model 113 (S503). As an update method, the backpropagation method in a neural network can be used.
[0040] Next, the learning program 411 determines whether the learning has converged (S504). The convergence determination is performed under conditions such as having executed a fixed number of times and the value of the error function falling below a defined threshold.
[0041] FIG. 6 is a diagram showing an example of a flowchart of the differential emotion recognition program 421.
[0042] Before the execution of the differential emotion recognition program 421, the voice to be analyzed by the user 102 etc. is stored in the storage device 400 as continuous voice data 121.
[0043] First, the differential emotion recognition program 421 reads in one frame of the continuous voice data 121, and if it has finished reading the continuous voice data 121, the program ends (S601).
[0044] Next, the differential emotion recognition program 421 determines whether a voice section has been detected (S602). For the detection of the voice section, known methods can be used. For example, it is a method where when a certain number of consecutive frames with a volume above a certain value are followed by a certain number of consecutive frames with a volume below a certain value, the group of frames up to that point is regarded as the voice section. If a voice section has not been detected, the process returns to S601. Note that the voice section may be the section where the voice is detected (a series of frame groups), or it may be a section including the frames before and / or after the section where the voice is detected.
[0045] Next, the differential emotion recognition program 421 stores the information of the detected voice section in the emotion transition data 122 (S603). Note that in S603, the voice section ID and the voice section data are stored in the emotion transition data 122, and the emotion transition is stored in S606.
[0046] Next, the differential emotion recognition program 421 determines whether it is possible to select a pair of voice segments (S604). The pair of voice segments to be selected can be, for example, a pair that is temporally adjacent in the emotion transition data 122, and can include a voice segment for which the emotion transition has not been calculated as one of the pair. Here, in order to make the calculation of the emotion transition robust, the differential emotion recognition program 421 may use adjacent pairs as a plurality of pairs within a predetermined time. Further, the differential emotion recognition program 421 may perform a process of excluding pairs including voice segments for which emotion recognition is considered difficult, such as short voice segments and voice segments with low volume.
[0047] Next, the differential emotion recognition program 421 inputs the selected pair of voice segments (voice data of two voice segments) into the differential emotion recognition model 113, and obtains the differential emotion output from the differential emotion recognition model 113 (S605).
[0048] Next, the differential emotion recognition program 421 calculates the emotion transition based on the differential emotion, and stores it in the emotion transition data 122 (S606). To obtain the emotion transition of a certain voice segment, for example, for all pairs, a value obtained by adding the obtained differential emotion to the emotion transition of the other voice segment paired with that voice segment is obtained, and the average value of that value is taken. The emotion transition of the first voice segment may be the average value "0". Thereafter, the process returns to S601.
[0049] In FIG. 6, a configuration in which a pair of voice segments detected from the continuous voice data 121 is input into the differential emotion recognition model 113 to obtain the differential emotion is taken as an example, but the configuration is not limited thereto. For example, the differential emotion recognition program 421 may be configured to input two pieces of voice data specified by the user into the differential emotion recognition model 113 to obtain the differential emotion.
[0050] FIG. 7 is a diagram showing an example (emotion transition table 700) of the emotion transition data 122.
[0051] The emotional transition table 700 stores emotional transitions for each voice segment. The voice segment ID is a code that uniquely identifies the voice segment. The voice segment data is information (e.g., time interval information) indicating from which position to which position in the continuous voice data 121 the voice segment is. The emotional transition is the emotional transition value of the voice segment obtained by the differential emotion recognition program 421.
[0052] FIG. 8 is a diagram showing an example of the user interface of the differential emotion recognition device 101.
[0053] The user 102 can obtain information from the display 402 that it is possible to select a continuous voice file to input. When the user 102 operates the keyboard 403 and / or the mouse 404 to press the voice file selection button 801 and selects a voice file stored in the differential emotion recognition device 101, the voice of the voice file is visualized as a waveform 810 on the display 402 and stored in the storage device 400 as continuous voice data 121. The user 102 can execute the differential emotion recognition program 421 by subsequently pressing the analysis start button 802. When the emotional transition data 122 is generated, the emotional transition value is visualized as a graph 820 on the display 402.
[0054] Note that since the emotional transition value is calculated for each voice segment, in the graph 820, the emotional transition values are smoothly connected and the emotional transition is shown in time series. Additionally, the data to be visualized is not limited to the emotional transition value, and may be the category of emotion or the degree (differential emotion value) for each emotion category.
[0055] If the emotion recognition system is configured with the content described above, the user can easily confirm the emotional transition of the voice of a specific speaker by the differential emotion recognizer using the model that has learned the differential value of the emotional expression of the voice.
[0056] (II) Second Embodiment In a conventional system designed to output an absolute emotion evaluation value (absolute value of emotion) for the input voice, there is a risk that the user may input voices of different people and use the emotion evaluation value to evaluate the emotional aspect of each person. As described above, since the output of the emotion recognizer reflects a small number of subjectivities and the accuracy is not sufficient, there are situations where such use is not appropriate. As the main use, it is often sufficient to be able to observe relative changes in emotion, such as visualizing the fluctuations in emotional expression for the voice of the same person, but no mechanism is provided to limit it to such uses.
[0057] In this regard, the differential emotion recognition device 901 of the present embodiment determines whether the two input voice data are voice data of the same speaker. In the present embodiment, for the same configuration as that of the first embodiment, the same reference numerals are used and the description thereof is omitted.
[0058] FIG. 9 is a diagram showing an example of a processing flow according to the differential emotion recognition device 901 of the present embodiment.
[0059] The differential emotion recognition model 911 in the present embodiment includes a same speaker recognition unit and has an output z in addition to the output y as an output layer. The output z is a same speaker determination value, and outputs "1" when the first voice and the second voice are of the same speaker, and "0" when they are not of the same speaker. The output z is calculated by (Equation 3).
Equation
[0060] In the learning program 411, when learning using the learning voice data 111 and the learning differential emotion label data 112 similar to those in the first embodiment, the label z becomes "1". Otherwise, two arbitrary voice data with different speaker IDs are extracted from the learning voice data 111, and the parameter is updated with the label y being a random value (a value unrelated to the differential emotion of the learning differential emotion label data 112) and the label z being "0".
[0061] When calculating the emotion transition in the differential emotion recognition program 421, if the same speaker determination value is less than the threshold, it is determined that they are not the same speaker. When the differential emotion value determined to be not the same speaker is a certain ratio or more of the total pairs, the emotion transition is set to an invalid value. In the visualization of the emotion transition by the display 402, it is displayed that the emotion recognition is invalid for the result of the voice section that is an invalid value. Additionally, the differential emotion recognition program 421 may abort the analysis when the emotion transition is set to an invalid value.
[0062] Note that this embodiment is not limited to a configuration in which the differential emotion recognition model 911 includes a same speaker recognition unit. For example, a configuration may be adopted in which the differential emotion recognition model 113 and a same speaker recognition unit (for example, a same speaker recognition model of a neural network) are provided in the differential emotion recognition device 901.
[0063] If the emotion recognition system is configured with the content described above, when the user attempts to perform emotion recognition including voices of different speakers, the result cannot be obtained. Thereby, it is possible to avoid a usage method in which the emotion evaluation value when voices of different people are input is regarded as an evaluation of the emotional aspect for each person.
[0064] (III) Supplementary Note The above-described embodiment includes, for example, the following content.
[0065] In the above-described embodiments, the case where the present invention is applied to an emotion recognition system has been described. However, the present invention is not limited thereto, and can be widely applied to various other systems, apparatuses, methods, and programs.
[0066] Also, in the above-described embodiments, the processing may be described with "program" as the subject. However, the program is executed by the processor unit to perform the defined processing while appropriately using a storage unit (e.g., memory) and / or an interface unit (e.g., communication port), etc. Therefore, the subject of the processing may be the processor. The processing described with the program as the subject may be the processing performed by the processor unit or an apparatus having the processor unit. Further, the processor unit may include a hardware circuit (e.g., FPGA (Field-Programmable Gate Array) or ASIC (Application Specific Integrated Circuit)) that performs part or all of the processing.
[0067] Also, in the above-described embodiments, part or all of the program may be installed from the program source into an apparatus such as a computer that realizes the differential emotion recognition device. The program source may be, for example, a program distribution server connected by a network or a computer-readable recording medium (e.g., a non-temporary recording medium). Further, in the above description, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0068] Also, in the above-described embodiments, the configuration of each table is an example, and one table may be divided into two or more tables, or all or part of two or more tables may be one table.
[0069] Also, in the above-described embodiments, the screens illustrated and described are examples, and any design may be used as long as the received information is the same.
[0070] Also, in the above-described embodiments, the screens illustrated and described are merely examples, and any design may be used as long as the information presented is the same.
[0071] Also, in the above-described embodiments, the case where the average value is used as the statistical value has been described. However, the statistical value is not limited to the average value, and other statistical values such as the maximum value, minimum value, difference between the maximum value and the minimum value, mode, median, standard deviation, etc. may also be used.
[0072] Also, in the above-described embodiments, the output of information is not limited to the display on the display. The output of information may be voice output by a speaker, output to a file, printing on a paper medium or the like by a printing device, projection on a screen or the like by a projector, or any other mode.
[0073] Also, in the above description, information such as programs, tables, and files that implement each function can be stored in a memory, a storage device such as a hard disk or an SSD (Solid State Drive), or a recording medium such as an IC card, an SD card, or a DVD.
[0074] The above-described embodiments have, for example, the following characteristic configurations.
[0075] (1) An emotion recognition system (e.g., differential emotion recognition device 101, differential emotion recognition device 901, a system including a differential emotion recognition device 101 and a computer capable of communicating with the differential emotion recognition device 101, a system including a differential emotion recognition device 901 and a computer capable of communicating with the differential emotion recognition device 901) includes an input unit (e.g., differential emotion recognition program 421, input unit 422, circuit) that inputs first voice data and second voice data, and inputs the first voice data and the second voice data to a differential emotion recognition model (e.g., differential emotion recognition model 113, differential emotion recognition model 911) that infers the differential emotion in the two voice data, and a processing unit (e.g., differential emotion recognition program 421, processing unit 423, circuit) that obtains differential emotion information indicating the differential emotion in the first voice data and the second voice data from the differential emotion recognition model. Note that the first voice data and the second voice data may be included in one voice data (e.g., continuous voice data 121) or may be separate voice data.
[0076] In the above configuration, for example, since the differential emotion (i.e., relative emotion) is inferred from two voice data, the emotion recognized by the emotion recognition system can be made closer to the impression of the actual user than the emotion inferred from one voice data (i.e., absolute emotion).
[0077] (2) The above-mentioned emotion recognition system inputs the first voice data and the second voice data into an identical speaker recognition unit (e.g., differential emotion recognition model 911, identical speaker recognition unit) that infers whether the speakers of the two voice data are the same speaker, and obtains determination information (e.g., it may be "0" indicating not the same speaker or "1" indicating the same speaker, or it may be a numerical value from "0" to "1") indicating that the speaker of the first voice data and the speaker of the second voice data are the same speaker from the identical speaker recognition unit, and a determination unit (e.g., differential emotion recognition program 421 including a determination unit, determination unit, circuit) that determines whether the speaker of the first voice data and the speaker of the second voice data are the same speaker according to the obtained determination information, and an output unit (e.g., differential emotion recognition program 421 including a determination unit, output unit 424, circuit) that outputs information corresponding to the result of the determination by the determination unit.
[0078] According to the above configuration, since it is determined whether the speakers of the two voice data are the same speaker, for example, incorrect usage of inputting voices of different people and evaluating the emotions of each person can be avoided.
[0079] (3) When it is determined by the determination unit that they are the same speaker, the output unit outputs the differential emotion information (e.g., graph 820) obtained by the processing unit, and when it is determined by the determination unit that they are not the same speaker, it outputs that it is not the voice of the same speaker (e.g., "There is a voice section of different speakers. The graph of that voice section is not displayed.").
[0080] According to the above configuration, for example, a mechanism can be provided to reject the comparison of voices between different persons.
[0081] (4) The above input unit inputs continuous voice data (for example, continuous voice data 121), and the above processing unit detects a voice section from the continuous voice data input by the above input unit, extracts voice data from the continuous voice data for each voice section, selects one voice data and other voice data within a predetermined time from the one voice data from the extracted voice data, inputs the one voice data and the other voice data into the above differential emotion recognition model, and obtains differential emotion information indicating the differential emotion between the one voice data and the other voice data (for example, refer to FIG. 6).
[0082] According to the above configuration, when continuous voice data is input, for example, differential emotion information indicating the differential emotion in two adjacent voice data is sequentially obtained, so that the user can grasp the transition of emotions.
[0083] (5) The above differential emotion recognition model is learned using two voice data of the same person and differential emotion information indicating the differential emotion in the two voice data (refer to FIG. 5).
[0084] In the above configuration, since two voice data of the same person are labeled, for example, it is easier for the labeler to estimate emotions, and the situation where the label depends on the subjectivity of the labeler can be reduced.
[0085] (6) When the above same speaker recognition unit is learned using two voice data, differential emotion information indicating the differential emotion in the two voice data, and information indicating whether the speakers of the two voice data are the same, when the two voice data are voice data of different persons, the differential emotion information indicating the differential emotion in the two voice data is changed to a value unrelated to the above differential emotion information (for example, a random value) and learned.
[0086] According to the above configuration, for example, the differential emotion recognition model and the same speaker recognition unit can be trained using common data, so that the burden of preparing the data used for training can be reduced.
[0087] (7) The above differential emotion recognition model is a neural network.
[0088] In the above configuration, since the differential emotion recognition model is a neural network, it is possible to reduce the situation where the rules do not function correctly due to influencing factors such as speaker characteristics such as people with a tendency to muffle their voices or people with a high voice, and environmental characteristics such as large reverberation, and improve the accuracy of inference.
[0089] Regarding the above-described configuration, within a range not exceeding the gist of the present invention, it may be appropriately changed, recombined, combined, or omitted.
[0090] It should be understood that the items included in the list in the form of "at least one of A, B, and C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). Similarly, the items listed in the form of "at least one of A, B, or C" can mean (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
Explanation of Signs
[0091] 101... Differential emotion recognition device, 113... Differential emotion recognition model.
Claims
1. An input unit that inputs first voice data and second voice data, A differential emotion recognition model that infers the differential emotion in two voice data inputs the first voice data and the second voice data, and a processing unit that obtains differential emotion information indicating the differential emotion in the first voice data and the second voice data from the differential emotion recognition model, Comprising, The differential emotion recognition model is trained using the two voice data of the same person and differential emotion information indicating the differential emotion in the two voice data. An emotion recognition system.
2. An identical speaker recognition unit that infers whether the speakers of the two voice data are the same inputs the first voice data and the second voice data, and obtains determination information indicating that the speaker of the first voice data and the speaker of the second voice data are the same from the identical speaker recognition unit, and a determination unit that determines whether the speaker of the first voice data and the speaker of the second voice data are the same according to the obtained determination information, An output unit that outputs information according to the result of the determination by the determination unit, The emotion recognition system according to claim 1, comprising.
3. When the output unit is determined to be the same speaker by the determination unit, it outputs the differential emotion information obtained by the processing unit, and when it is determined by the determination unit not to be the same speaker, it outputs that it is not the voice of the same speaker. The emotion recognition system according to claim 2.
4. The input unit inputs continuous voice data, The processing unit detects a voice section from the continuous voice data input by the input unit, extracts voice data from the continuous voice data for each voice section, selects one voice data and other voice data within a predetermined time from the extracted voice data, inputs the one voice data and the other voice data into the differential emotion recognition model, and obtains differential emotion information indicating the differential emotion in the one voice data and the other voice data. The emotion recognition system according to claim 1.
5. When the same speaker recognition unit is trained using the two voice data, difference emotion information indicating the difference in emotion between the two voice data, and information indicating whether the speakers of the two voice data are the same, when the two voice data are voice data of different persons, the difference emotion information indicating the difference in emotion between the two voice data is changed to a value independent of the difference emotion information indicating the difference in emotion between the first voice data and the second voice data and is trained. The emotion recognition system according to claim 2.
6. The difference emotion recognition model is a neural network. The emotion recognition system according to claim 1.
7. An input unit inputs first voice data and second voice data. A processing unit inputs the first voice data and the second voice data to a difference emotion recognition model that infers the difference in emotion between two voice data, and obtains difference emotion information indicating the difference in emotion between the first voice data and the second voice data from the difference emotion recognition model. including The difference emotion recognition model is trained using the two voice data of the same person and difference emotion information indicating the difference in emotion between the two voice data. Emotion recognition method.
8. A determination unit inputs the first voice data and the second voice data to an identical speaker recognition unit that infers whether the speakers of the two voice data are the same speaker, and obtains determination information indicating that the speaker of the first voice data and the speaker of the second voice data are the same speaker from the identical speaker recognition unit, and determines whether the speaker of the first voice data and the speaker of the second voice data are the same speaker according to the obtained determination information. An output unit outputs information according to the result of the determination by the determination unit. The emotion recognition method according to claim 7, including
9. When the output unit is determined to be the same speaker by the determination unit, the difference emotion information obtained by the processing unit is output, and when the output unit is determined not to be the same speaker by the determination unit, it outputs that it is not the voice of the same speaker. The emotion recognition method according to claim 8.
10. The input unit inputs continuous voice data. The processing unit detects a voice section from the continuous voice data input by the input unit, extracts voice data from the continuous voice data for each voice section, selects one voice data and other voice data within a predetermined time from the one voice data from the extracted voice data, inputs the one voice data and the other voice data into the differential emotion recognition model, and obtains differential emotion information indicating the differential emotion in the one voice data and the other voice data. The emotion recognition method according to claim 7.
11. When the same speaker recognition unit is learned using the two voice data, the differential emotion information indicating the differential emotion in the two voice data, and the information indicating whether the speakers of the two voice data are the same, when the two voice data are voice data of different persons, the differential emotion information indicating the differential emotion in the two voice data is changed to a value unrelated to the differential emotion information indicating the differential emotion in the first voice data and the second voice data and is learned. The emotion recognition method according to claim 8.
12. The differential emotion recognition model is a neural network. The emotion recognition method according to claim 7.
Citation Information
Patent Citations
Voice interaction apparatus
JP2018132623A
Emotion estimation system and program
JP2020008730A
Information processing device, information processing method, recognition model and program
JP2021026130A
Information processing device
WO2017187712A1
Voice processing device, voice processing method, and program storage medium
WO2020049687A1