Speaker identification device, speaker identification method, and program
The speaker identification device uses a DNN to estimate emotions in speech and adapt speaker identification processing, improving accuracy by matching emotional states in registered and evaluation utterances.
Patent Information
- Application Number
- JP2022503218
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2021-02-05
- Publication Date
- 2025-09-11
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Conventional speaker identification techniques fail to accurately identify speakers when emotional speech such as laughter or yelling is used due to intonation fluctuations caused by differing emotions between registered and evaluation speech.
A speaker identification device utilizing a trained Deep Neural Network (DNN) to estimate emotions in speech and adjust speaker identification processing based on the estimated emotions, incorporating emotion-specific speaker classifiers or feature corrections to improve accuracy.
Enhances speaker identification accuracy by matching the emotions in registered and evaluation utterances, enabling precise identification even in the presence of emotional speech.
Smart Images

Figure 0007737976000001 
Figure 0007737976000002 
Figure 0007737976000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a speaker identification device, a speaker identification method, and a program. [Background technology]
[0002] Speaker identification technology is a technology that estimates which speaker's registered utterance an evaluation utterance is based on the similarity between features calculated from registered utterances, which are utterances for each speaker to be registered, and features calculated from evaluation utterances, which are utterances for an unknown speaker to be identified (for example, Patent Document 1).
[0003] For example, Patent Document 1 discloses a speaker identification technology that identifies the speaker of an evaluation utterance by using the vector similarity between a speaker feature vector in a registration utterance for each registration speaker and a speaker feature vector in an evaluation utterance. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2017-187642 Summary of the Invention [Problem to be solved by the invention]
[0005] However, if emotional speech such as laughter or yelling is used as the evaluation speech, it will affect the accuracy of the identification. Specifically, if the emotion contained in the registered speech differs from the emotion contained in the evaluation speech, the accuracy of speaker identification will decrease due to the intonation fluctuations associated with the emotion contained in the evaluation speech.
[0006] In other words, conventional speaker identification techniques such as that disclosed in Patent Document 1 identify the speaker of an evaluation utterance by calculating the similarity between the speaker feature vectors of a registered utterance and an evaluation utterance without taking into account the emotion contained in the evaluation utterance. For this reason, conventional speaker identification techniques may not be able to identify the speaker of an evaluation utterance with sufficient accuracy.
[0007] The present disclosure has been made in consideration of the above-mentioned circumstances, and aims to provide a speaker identification device, a speaker identification method, and a program that can improve the accuracy of speaker identification even when the evaluation utterance, i.e., the utterance to be identified, contains the speaker's emotions. [Means for solving the problem]
[0008] A speaker identification device according to one aspect of the present disclosure is a speaker identification device that identifies a speaker of speech data indicating the speech of an utterance to be identified, and includes: an emotion estimator that uses a trained DNN (Deep Neural Network) to estimate an emotion included in the speech of the utterance indicated by the speech data from acoustic features calculated from the utterance data; and a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying the speaker of the speech data from the acoustic features calculated from the utterance data.
[0009] These general or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]
[0010] According to the speaker identification device and the like of the present disclosure, it is possible to improve the accuracy of speaker identification even when the speech to be identified includes the emotion of the speaker. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a speaker identification system according to an embodiment. [Figure 2] FIG. 2 is a block diagram showing another example of the configuration of the speaker identification system according to the embodiment. [Figure 3] FIG. 3 is a block diagram illustrating an example of a detailed configuration of the preprocessing unit according to the embodiment. [Figure 4] FIG. 4 is a block diagram illustrating an example of a detailed configuration of a speaker identification device according to an embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a configuration of the emotion estimator according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of the configuration of a speaker identifier according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of the configuration of a speaker feature extraction unit included in a speaker classifier according to an embodiment. [Figure 8] FIG. 8 is a flowchart showing an outline of the operation of the speaker identification device according to the embodiment. [Figure 9] FIG. 9 is a block diagram showing an example of a detailed configuration of a speaker identification device according to the first modification of the embodiment. [Figure 10] FIG. 10 is a block diagram showing an example of a detailed configuration of a speaker identification device according to the second modification of the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of processing performed by the speaker identification device according to the second modification of the embodiment. [Figure 12] FIG. 12 is a block diagram showing an example of a detailed configuration of a speaker identification device according to the third modification of the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] (Summary of the Disclosure) An outline of one embodiment of the present disclosure is as follows.
[0013] A speaker identification device according to one aspect of the present disclosure is a speaker identification device that identifies a speaker of speech data indicating the speech of an utterance to be identified, and includes: an emotion estimator that uses a trained DNN (Deep Neural Network) to estimate an emotion included in the speech of the utterance indicated by the speech data from acoustic features calculated from the utterance data; and a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying the speaker of the speech data from the acoustic features calculated from the utterance data.
[0014] According to this aspect, it is possible to improve the accuracy of speaker identification even if the speech to be identified contains the emotion of the speaker.
[0015] Further, for example, the speaker identification processing unit may include a plurality of speaker classifiers each having: a speaker feature extraction unit that, when the acoustic features are input, extracts first speaker features that can identify a speaker of the speech of the utterance indicated by the speech data from the input acoustic features; and a similarity calculation unit that calculates a similarity between the first speaker features extracted by the speaker feature extraction unit and second speaker features that are stored in a storage unit and can identify each speech containing one emotion of the registered speaker to be identified; and a classifier selection unit that selects one of the plurality of speaker classifiers, the speaker classifier having stored in the storage unit second speaker features that can identify each speech containing one emotion of the registered speaker corresponding to the emotion indicated by the estimation result, and the speaker classifier selected by the classifier selection unit may calculate the similarity when the acoustic features calculated from the utterance data are input, and output the similarity as the score.
[0016] Furthermore, for example, the speaker identification processing unit may include a speaker feature extraction unit that extracts, from the acoustic features, first speaker features that can identify the speaker of the speech indicated by the speech data; a correction unit that corrects second speaker features that are stored in a storage unit and can identify each speech containing one emotion of a registered speaker to be identified, into third speaker features that can identify each speech containing the one emotion corresponding to the emotion indicated by the estimation result; and a similarity calculation unit that calculates a similarity between the extracted first speaker features and the third speaker features corrected by the correction unit, and outputs the calculated similarity as the score.
[0017] Furthermore, for example, the speaker identification processing unit may include a speaker feature extraction unit that extracts, from the acoustic features, first speaker features that can identify the speaker of the voice of the utterance indicated by the speech data; a similarity calculation unit that calculates a similarity between the extracted first speaker feature and a second speaker feature that is stored in a storage unit and can identify each voice containing one emotion of the registered speaker to be identified; and a reliability assignment unit that assigns a weighting to the calculated similarity in accordance with the emotion indicated by the estimation result and outputs the calculated similarity as the score, and the reliability assignment unit may assign the largest weighting to the calculated similarity when the one emotion matches the emotion indicated by the estimation result.
[0018] Furthermore, for example, the acoustic feature may be calculated from each of a plurality of pieces of speech data acquired by a preprocessing unit by dividing entire speech data indicating speech sounds of one speaker over a predetermined period into identification units in a time series, and the reliability assigning unit may assign a weight to the similarity for each of the plurality of pieces of speech data calculated by the similarity calculation unit in accordance with an emotion indicated by the estimation result for each of the plurality of pieces of speech data estimated by the emotion estimator, and output the weight as the score.
[0019] Also, for example, the speaker identification device may further include a speaker identification unit that identifies the speaker of the entire speech data using an overall score, which is an arithmetic average of the scores for each of the multiple speech data output by the reliability assignment unit, and the speaker identification unit may identify the speaker of the entire utterance using an overall score that is equal to or greater than a threshold value.
[0020] Furthermore, for example, the speaker identification processing unit may include a speaker feature extraction unit that extracts, from the acoustic features, first speaker features that can identify the speaker of the voice of the utterance indicated by the speech data; a similarity calculation unit that calculates a similarity between the extracted first speaker features and second speaker features that are stored in a storage unit and can identify each voice containing one emotion of the registered speaker to be identified; and a reliability assignment unit that assigns a reliability to the calculated similarity according to the emotion indicated by the estimation result and outputs the calculated similarity as the score.
[0021] Furthermore, for example, the speaker identification device may further include a speaker identification unit that identifies the speaker of the utterance data by using the score where the reliability is equal to or greater than a threshold value.
[0022] Furthermore, for example, the speaker feature extraction unit may extract the first speaker feature from the acoustic feature using a trained DNN.
[0023] A speaker identification method according to one embodiment of the present disclosure is a speaker identification method for identifying a speaker of speech data indicating the speech of an utterance to be identified, and includes an emotion estimation step of estimating, using a trained DNN, an emotion included in the speech indicated by the speech data from acoustic features calculated from the utterance data, and a speaker identification processing step of outputting, using an estimation result in the emotion estimation step, a score for identifying the speaker of the speech data from the acoustic features calculated from the utterance data.
[0024] Furthermore, a program according to one aspect of the present disclosure is a program that causes a computer to execute a speaker identification method for identifying a speaker of speech data indicating the speech of an utterance to be identified, and causes the computer to execute an emotion estimation step of estimating an emotion contained in the speech indicated by the speech data from acoustic features calculated from the utterance data using a trained DNN, and a speaker identification processing step of outputting a score for identifying the speaker of the speech data from the acoustic features calculated from the utterance data using an estimation result in the emotion estimation step.
[0025] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.
[0026] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Each of the embodiments described below illustrates a specific example of the present disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in an independent claim that represents a top-level concept are described as optional components. Furthermore, in all embodiments, the respective contents can be combined.
[0027] (Embodiment) Hereinafter, a speaker identification device according to the present embodiment will be described with reference to the drawings.
[0028] [Speaker Identification System 1] Fig. 1 is a block diagram showing an example of the configuration of a speaker identification system 1 according to this embodiment. Fig. 2 is a block diagram showing another example of the configuration of a speaker identification system 1 according to this embodiment.
[0029] The speaker identification system 1 according to this embodiment is used to identify the speaker of utterance data indicating the voice of an utterance to be identified, which utterance includes the emotion of the speaker.
[0030] As shown in Fig. 1, the speaker identification system 1 includes a preprocessing unit 10 and a speaker identification device 11. Note that the speaker identification system 1 may further include a speaker identification unit 14 as shown in Fig. 2, but this is not an essential component. Each component will be described below.
[0031] [1. Pre-processing unit 10] FIG. 3 is a block diagram showing an example of a detailed configuration of the preprocessing unit 10 according to this embodiment.
[0032] The preprocessing unit 10 acquires speech data indicating the speech of the speech to be identified, and outputs acoustic features calculated from the acquired speech data to the speaker identification device 11. In this embodiment, the preprocessing unit 10 includes a speech acquisition unit 101 and an acoustic feature calculation unit 102, as shown in FIG.
[0033] [1.1 Voice Acquisition Unit 101] The speech acquisition unit 101 is, for example, a microphone, and acquires the speech of a speaker. The speech acquisition unit 101 converts the acquired speech into a speech signal, detects a speech section where speech is made, and outputs speech data indicating the speech obtained by extracting the speech section to the acoustic feature calculation unit 102.
[0034] The speech acquisition unit 101 may acquire multiple pieces of speech data by dividing the entire speech data indicating the speech of one speaker over a predetermined period into identification units in a time series, and output the pieces of speech data to the acoustic feature calculation unit 102. The identification unit may be, for example, 3 to 4 seconds, and may be the above-mentioned speech section.
[0035] [1.2 Acoustic feature calculation unit 102] The acoustic feature calculation unit 102 calculates acoustic features for the speech of the utterance from the speech signal of the speech section output by the speech acquisition unit 101, i.e., the speech data. In this embodiment, the acoustic feature calculation unit 102 calculates MFCCs (Mel Frequency Cepstral Coefficients), which are features of the speech of the utterance, as the acoustic features from the speech data output by the speech acquisition unit 101. MFCCs are features that represent the vocal tract characteristics of a speaker and are also commonly used in speech recognition. More specifically, MFCCs are acoustic features obtained by analyzing the frequency spectrum of speech based on the characteristics of human hearing. Note that the acoustic feature calculation unit 102 is not limited to calculating MFCCs as acoustic features from the speech data. Alternatively, the acoustic feature calculation unit 102 may calculate, as the acoustic feature, a result obtained by applying a Mel filter bank to the speech speech signal, or may calculate, as the acoustic feature, a spectrogram of the speech speech signal.
[0036] [2. Speaker Identification Device 11] The speaker identification device 11 is realized by, for example, a computer including a processor (microprocessor), a memory, a communication interface, etc. The speaker identification device 11 may be included in a server and operate, or a portion of the configuration of the speaker identification device 11 may be included in a cloud server and operate. The speaker identification device 11 performs processing to identify the speaker of utterance data indicating the voice of an evaluation utterance, i.e., an utterance to be identified. More specifically, the speaker identification device 11 outputs, as an identification result, a score representing the similarity between a first speaker feature of the evaluation utterance and a second speaker feature of a registered utterance for each registered speaker. The evaluation utterance, i.e., the utterance to be identified, according to this embodiment includes the emotion of the speaker.
[0037] FIG. 4 is a block diagram showing an example of a detailed configuration of the speaker identification device 11 according to this embodiment.
[0038] The speaker identification device 11 includes an emotion estimator 12 and a speaker identification processing unit 13, as shown in FIGS.
[0039] [2.1 Emotion Estimator 12] The emotion estimator 12 uses a trained DNN (Deep Neural Network) to estimate the emotion contained in the speech represented by the speech data from acoustic features calculated from the speech data. Note that the DNN may be, for example, a CNN (Convolution Neural Network), a fully connected NN (Neural Network), or a TDNN (Time Delay Neural Network).
[0040] An example of the configuration of the emotion estimator 12 will now be described with reference to FIG.
[0041] FIG. 5 is a diagram showing an example of the configuration of the emotion estimator 12 according to the present embodiment.
[0042] The emotion estimator 12 includes a frame connection processing unit 121 and a DNN 122, as shown in FIG.
[0043] [2.1.1 Frame connection processing unit 121] The frame connection processing unit 121 connects multiple frames of MFCCs, which are acoustic features output from the preprocessing unit 10, and outputs the result to the input layer of the DNN 122. The MFCCs are configured with multiple frames, each having an x-dimensional feature (x is a positive integer) for one frame. In the example shown in FIG. 5, the frame connection processing unit 121 connects 50 frames of MFCC parameters, each consisting of 24-dimensional / frame features, to generate a 1200-dimensional vector and outputs the vector to the input layer of the DNN 122.
[0044] [2.1.2 DNN122] When multiple frames of the connected MFCCs are input, the DNN 122 outputs the emotion label with the highest probability as the estimation result of the emotion estimator 12. In the example shown in FIG. 5, the DNN 122 is a neural network consisting of an input layer, multiple hidden layers, and an output layer, and is trained using training data stored in the memory unit 123, i.e., training speech data containing the emotion to be estimated. The input layer consists of, for example, 1200 nodes, and receives as input a 1200-dimensional vector generated by connecting 50 frames of MFCC parameters consisting of features of 24 dimensions per frame. The output layer consists of nodes that output emotion labels such as neutral, anger, laughter, and sadness, and outputs the emotion label with the highest probability. The multiple hidden layers consist of, for example, two to three hidden layers.
[0045] [2.2 Speaker identification processing unit 13] The speaker identification processing unit 13 uses the estimation result of the emotion estimator 12 to output a score for identifying the speaker of the utterance data from the acoustic feature amount calculated from the utterance data.
[0046] In this embodiment, the speaker identification processing unit 13 includes a classifier selection unit 131 and a plurality of speaker classifiers 132, as shown in FIG.
[0047] 2.2.1 Multiple Speaker Classifiers 132 Each of the multiple speaker classifiers 132 is a speaker classifier 132k (k is a natural number) corresponding to one emotion. An emotion is, for example, one of calm, anger, laughter, sadness, and so on. In the example shown in FIG. 4, the multiple speaker classifiers 132 are made up of speaker classifier 132a, speaker classifier 132b, and so on. For example, speaker classifier 132a corresponds to one emotion, calm, and speaker classifier 132b corresponds to one emotion, laughter. Note that one of speaker classifier 132a, speaker classifier 132b, and so on is represented as speaker classifier 132k.
[0048] The speaker classifier 132k selected by the classifier selection unit 131 from among the plurality of speaker classifiers 132 receives the acoustic features calculated from the speech data, calculates the similarity, and outputs it as a score. Note that there are cases where none of the plurality of speaker classifiers 132 is selected by the classifier selection unit 131, and in Fig. 4, this is expressed as a case where the classifier selection unit 131 selects "no selection."
[0049] Here, as an example of the configuration of the speaker classifier 132k, the speaker classifier 132b corresponding to laughter will be described with reference to FIG.
[0050] Fig. 6 is a diagram showing an example of the configuration of speaker classifier 132b according to this embodiment. Fig. 7 is a diagram showing an example of the configuration of speaker feature extraction unit 133b included in speaker classifier 132b according to this embodiment.
[0051] As shown in FIG. 6, the speaker identifier 132b includes a speaker feature extractor 133b, a storage unit 134b, and a similarity calculator 135b.
[0052] [2.2.1.1 Speaker feature extraction unit 133b] When acoustic features calculated from speech data are input, the speaker feature extraction unit 133b extracts, from the input acoustic features, first speaker features that can identify the speaker of the speech represented by the speech data. More specifically, the speaker feature extraction unit 133b extracts the first speaker features from the acoustic features using a trained DNN.
[0053] In this embodiment, the speaker feature extraction unit 133b extracts the first speaker feature using, for example, the x-vector method. Here, the x-vector method is a method for calculating a speaker feature, which is a feature unique to a speaker, called an x-vector. More specifically, the speaker feature extraction unit 133b includes, for example, a frame connection processing unit 1331 and a DNN 1332b, as shown in FIG. 7.
[0054] [2.2.1.1-1 Frame connection processing unit 1331] The frame connection processing unit 1331 performs the same processing as the frame connection processing unit 121. That is, the frame connection processing unit 1331 connects multiple frames of MFCCs, which are acoustic features output from the preprocessing unit 10, and outputs the result to the input layer of the DNN 1332b. In the example shown in Fig. 7, the frame connection processing unit 1331 connects 50 frames of MFCC parameters, each consisting of features of 24 dimensions / frame, to generate a 1200-dimensional vector and outputs the vector to the input layer of the DNN 1332b.
[0055] [2.2.1.1-2 DNN1332b] The DNN 1332b outputs first speaker features when multiple frames are input from the frame connection processing unit 1331. In the example shown in Fig. 7, the DNN 1332b is a neural network consisting of an input layer, multiple intermediate layers, and an output layer, and is trained using training speech data, which is training data stored in the storage unit 1333b. In the example shown in Fig. 7, the storage unit 1333b stores training speech data composed of speeches of multiple speakers each containing laughter as one emotion.
[0056] In the example shown in FIG. 7, the input layer is composed of, for example, 1200 nodes, and receives as input a 1200-dimensional vector generated by connecting 50 frames of MFCC parameters each consisting of 24-dimensional / frame features. The output layer is composed of nodes that output speaker labels for the number of speakers included in the training data. The multiple intermediate layers are composed of, for example, two or three intermediate layers, and include an intermediate layer that calculates first speaker features. The intermediate layer that calculates the first speaker features outputs the calculated first speaker features as the output of the DNN 1332b.
[0057] [2.2.1.2 Storage unit 134b] The storage unit 134b is configured with, for example, a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores second speaker features that are unique features of pre-registered registered speakers and are calculated from registered utterances of the registered speakers. In other words, the storage unit 134b stores second speaker features that can identify each speech containing one emotion of the registered speaker. More specifically, the storage unit 134b stores second speaker features of registered utterances containing the emotion of laughter of the registered speaker, as shown in FIG. 6.
[0058] [2.2.1.3 Similarity calculation unit 135b] The similarity calculation unit 135b calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit 133b and the second speaker feature registered in advance and stored in the storage unit 134b.
[0059] In this embodiment, similarity calculation unit 135b calculates the similarity between the first speaker feature extracted by speaker feature extraction unit 133b and each of the second speaker features of one or more registered speakers stored in storage unit 134b. Similarity calculation unit 135b outputs a score representing the calculated similarity.
[0060] For example, the similarity calculation unit 135b may calculate the cosine distance (also referred to as cosine similarity) indicating the angle between the vectors of the first speaker feature amount and the second speaker feature amount by calculating the cosine using the dot product in the vector space model. In this case, a larger numerical value of the angle between the vectors indicates a lower similarity. Note that the similarity calculation unit 135b may also calculate, as the similarity, a cosine distance that takes a value between -1 and 1 using the dot product of the vector indicating the first speaker feature amount and the vector indicating the second speaker feature amount. In this case, a larger numerical value indicating the cosine distance indicates a higher similarity.
[0061] The speaker classifier 132a etc. corresponding to calmness is similar to the speaker classifier 132b corresponding to laughter, and therefore a description thereof will be omitted.
[0062] [2.2.2 Classifier selection unit 131] The classifier selection unit 131 selects one speaker classifier 132k from the plurality of speaker classifiers 132 according to the emotion indicated by the estimation result of the emotion estimator 12. More specifically, the classifier selection unit 131 selects the speaker classifier 132k that stores in a storage unit second speaker features that can identify each speech containing one emotion of a registered speaker according to the emotion indicated by the estimation result of the emotion estimator 12. Note that if there is no speaker classifier 132 that corresponds to the emotion indicated by the estimation result of the emotion estimator 12, the classifier selection unit 131 may not use any speaker classifier 132 (no selection).
[0063] In this way, the classifier selection unit 131 can switch the speaker classifier 132 depending on the estimation result of the emotion estimator 12.
[0064] [3. Speaker Identification Unit 14] When the speaker identification unit 14 is provided in the speaker identification system 1 as shown in FIG. 2, for example, it uses the score output by the speaker identification device 11 to identify the speaker of the speech data.
[0065] In this embodiment, the speaker identification unit 14 identifies the speaker of the utterance data based on the score representing the similarity calculated by the similarity calculation unit 135b. For example, by using such a score, the speaker identification unit 14 outputs, as an identification result, the registered speaker corresponding to the second speaker feature that is considered to be closest to the first speaker feature.
[0066] [Operation of speaker identification system 1] Next, a description will be given of the operation of the speaker identification system 1 configured as above. In the following, the operation of the speaker identification system 1, which is a characteristic operation of the speaker identification device 11, will be described.
[0067] FIG. 8 is a flowchart showing an outline of the operation of the speaker identification device 11 according to this embodiment.
[0068] First, the speaker identification device 11 estimates the emotion included in the voice of the utterance indicated by the utterance data from the acoustic feature calculated from the utterance data using the trained DNN (S11).
[0069] Next, the speaker identification device 11 uses the estimation result obtained in step S11 to output a score for identifying the speaker of the utterance data from the acoustic feature amount calculated from the utterance data (S12).
[0070] [Effects, etc.] As described above, according to the speaker identification device 11 of this embodiment, the emotion estimator 12 that estimates the emotion of an evaluation utterance is arranged in front of a plurality of speaker classifiers 132 each corresponding to one emotion, and the speaker classifiers 132 are switched depending on the emotion indicated in the estimation result of the emotion estimator 12.
[0071] This allows the use of a speaker identifier 132 that corresponds to the emotion of the evaluation utterance, so that the speaker of the evaluation utterance can be identified in a state where the emotion contained in the registered utterance matches the emotion contained in the evaluation utterance.
[0072] Therefore, the speaker identification device 11 according to this embodiment can improve the accuracy of speaker identification even if the speech to be identified contains the emotion of the speaker.
[0073] Furthermore, according to the speaker identification system 1 equipped with the speaker identification device 11 of this embodiment, it is possible to identify the speaker of an utterance, such as a conversation that is not a free speech, i.e., a reading of a sentence, in a meeting minutes system, a communication visualization system, or the like.
[0074] (Variation 1) Note that the method for identifying the speaker of utterance data indicating the audio of an utterance to be identified that includes the emotion of the speaker is not limited to the method described in the above embodiment, i.e., the method of configuring a plurality of speaker classifiers 132 subsequent to the emotion estimator 12. Below, an example of a method different from the method described in the above embodiment will be described as Variation 1, focusing on the differences from the above embodiment.
[0075] [4. Speaker Identification Device 11A] 9 is a block diagram showing an example of a detailed configuration of a speaker identification device 11A according to Modification 1 of the present embodiment. Note that the same elements as those in FIG. 4 and the like are given the same reference numerals, and detailed description thereof will be omitted.
[0076] The speaker identification device 11A performs processing to identify the speaker of utterance data indicating the voice of the utterance to be identified. More specifically, the speaker identification device 11A outputs, as an identification result, a score representing the similarity between a first speaker feature of the evaluation utterance and a third speaker feature obtained by correcting a second speaker feature of the registered utterance for each registered speaker.
[0077] As shown in FIG. 9, a speaker identification device 11A according to this modification differs from the speaker identification device 11 shown in FIG. 4 in the configuration of a speaker identification processing unit 13A.
[0078] [4.1 Speaker identification processing unit 13A] The speaker identification processing unit 13A uses the estimation result of the emotion estimator 12 to output a score for identifying the speaker of the utterance data from the acoustic feature amount calculated from the utterance data.
[0079] In this modification, as shown in FIG. 9, the speaker identification processing unit 13A includes a speaker feature extraction unit 133A, a storage unit 134A, a similarity calculation unit 135A, a storage unit 136A, and a correction unit 137A.
[0080] [4.1.1 Speaker feature extraction unit 133A] The speaker feature extracting unit 133A extracts, from the acoustic feature calculated from the speech data, a first speaker feature that can identify the speaker of the voice of the utterance indicated by the speech data.
[0081] In this modification, the speaker feature extraction unit 133A also extracts the first speaker feature using, for example, the x-vector method. Therefore, like the speaker feature extraction unit 133b, the speaker feature extraction unit 133A may include a frame connection processing unit and a DNN. In this modification, training is performed using training speech data consisting of the speech of each of multiple speakers to be classified, which includes, for example, calm as one emotion. Note that calm is an example of one emotion, and other emotions such as laughter may also be used. As the rest are as explained in the above embodiment, explanation here will be omitted.
[0082] 4.1.2 Storage Unit 134A The storage unit 134A is configured with, for example, a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores pre-registered second speaker features that can identify each speech containing one emotion of the registered speaker. In this modification, the storage unit 134A stores second speaker features of registered utterances containing the emotion of neutrality of the registered speaker, as shown in Fig. 9. Note that the emotion of neutrality is just one example, and other emotions such as laughter may also be used.
[0083] [4.1.3 Storage section 136A] The storage unit 136A is configured with a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores learning data for correcting emotions included in registered utterances. In this modification, the learning data stored in the storage unit 136A is used to correct the second speaker feature for the neutral emotion stored in the storage unit 134A to a third speaker feature, which is a speaker feature of an utterance of an emotion corresponding to the emotion indicated by the estimation result of the emotion estimator 12.
[0084] [4.1.4 Correction section 137A] The correction unit 137A corrects the second speaker feature stored in the storage unit 134A to a third speaker feature that can identify each speech containing one emotion corresponding to the emotion indicated by the estimation result of the emotion estimator 12.
[0085] For example, assume that the emotion indicated by the estimation result of the emotion estimator 12 is "laughter." In this case, the correction unit 137A uses the learning data stored in the storage unit 136A to correct the second speaker feature of a registered utterance containing the emotion "calm" of the registered speaker stored in the storage unit 134A to a third speaker feature that can identify each speech containing the emotion "laughter." In other words, the correction unit 137A uses the learning data stored in the storage unit 136A to correct the second speaker feature for the emotion "calm" stored in the storage unit 134A to a third speaker feature for the emotion indicated by the estimation result of the emotion estimator 12.
[0086] [4.1.5 Similarity calculation unit 135A] The similarity calculation unit 135A calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit 133A and the third speaker feature corrected by the correction unit 137A, and outputs the calculated similarity as a score.
[0087] In this modification, the similarity calculation unit 135A calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit 133A and each of the third speaker features obtained by correcting the second speaker feature of one or more registered speakers stored in the storage unit 134A. The similarity calculation unit 135A outputs a score representing the calculated similarity.
[0088] [5. Speaker Identification Unit 14] The speaker identification unit 14 uses the score output by the speaker identification device 11A to identify the speaker of the speech data.
[0089] In this modification, the speaker identification unit 14 identifies the speaker of the utterance data based on the score indicated by the similarity calculated by the similarity calculation unit 135A. For example, the speaker identification unit 14 uses the score to output, as an identification result, the registered speaker of the second speaker feature corresponding to the third speaker feature that is considered to be closest to the first speaker feature.
[0090] [Effects, etc.] As described above, according to the speaker identification device 11A of this modified example, the speaker identification processing unit 13A arranged in the subsequent stage corrects the emotion of the registered utterance to the emotion of the evaluation utterance according to the estimation result of the emotion estimator 12 arranged in the previous stage, and then identifies the speaker of the evaluation utterance.
[0091] This makes it possible to identify the speaker of the evaluation utterance while matching the emotion contained in the registered utterance with the emotion contained in the evaluation utterance, i.e., while correcting the difference in emotion, i.e., intonation, between the registered utterance and the evaluation utterance to match them.
[0092] Therefore, the speaker identification device 11A according to this modification can improve the accuracy of speaker identification even if the speech to be identified contains the emotion of the speaker.
[0093] (Variation 2) The method described in the above embodiment is not limited to the case described in the embodiment and modification 1. Below, a case where the speaker identification device has a different configuration from that described in the embodiment and modification 1 will be described.
[0094] [6. Speaker Identification Device 11B] Fig. 10 is a block diagram showing an example of a detailed configuration of a speaker identification device 11B according to Modification 2 of the present embodiment. Note that the same elements as those in Figs. 4 and 9 are given the same reference numerals, and detailed description thereof will be omitted.
[0095] Similar to the speaker identification device 11, the speaker identification device 11B performs processing to identify the speaker of utterance data indicating the voice of the utterance to be identified. More specifically, the speaker identification device 11B calculates the similarity between the first speaker feature of the evaluation utterance and the second speaker feature of the registered utterance for each registered speaker. Then, the speaker identification device 11B outputs the score obtained by assigning reliability to the calculated similarity as the identification result. In this modification, a case where weighting is assigned as reliability will be described.
[0096] As shown in Fig. 10, the speaker identification device 11B according to this modification has a different configuration of the speaker identification processing unit 13B from the speaker identification device 11 shown in Fig. 4. Also, the speaker identification device 11B according to this modification has a different configuration of the speaker identification processing unit 13B from the speaker identification device 11A shown in Fig. 9.
[0097] [6.1 Speaker identification processing unit 13B] The speaker identification processing unit 13B uses the estimation result of the emotion estimator 12 to output a score for identifying the speaker of the utterance data from the acoustic feature amount calculated from the utterance data.
[0098] Here, the acoustic features acquired by the speaker identification processing unit 13B are calculated from each of the multiple pieces of speech data obtained by the pre-processing unit 10 by dividing the entire speech data representing the speech of one speaker over a predetermined period into identification units in a time series.
[0099] In this modification, as shown in FIG. 10, the speaker identification processing unit 13B includes a speaker feature extraction unit 133A, a storage unit 134A, a similarity calculation unit 135B, and a reliability assignment unit 138B.
[0100] [6.1.1 Similarity calculation unit 135B] The similarity calculation unit 135B calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit 133A and the second speaker feature that is pre-registered and stored in the storage unit 134A and can identify each voice containing one emotion of the registered speaker to be identified.
[0101] In this modified example, the similarity calculation unit 135B calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit 133A and the second speaker feature in a registered utterance containing the emotion of “calm” of one or more registered speakers stored in the memory unit 134A.
[0102] [6.1.2 Trust Assignment Unit 138B] The confidence assigning unit 138B assigns a weighting to the similarity calculated by the similarity calculation unit 135B according to the emotion indicated by the estimation result of the emotion estimator 12, and outputs the result as a score. Here, when one emotion matches the emotion indicated by the estimation result, the confidence assigning unit 138B assigns the largest weighting to the calculated similarity.
[0103] In this modification, the reliability assigning unit 138B assigns weighting to the similarity for each of the plurality of utterance data calculated by the similarity calculation unit 135B according to the emotion indicated by the estimation result for each of the plurality of utterance data estimated by the emotion estimator 12. The reliability assigning unit 138B outputs the weighted similarity for each of the plurality of utterance data to the speaker identification unit 14 as a score for each of the plurality of utterance data.
[0104] [7. Speaker Identification Unit 14] When the speaker identification unit 14 is provided in the speaker identification system 1 as shown in FIG. 2, for example, it identifies the speaker of the speech data using the score output by the speaker identification device 11B.
[0105] In this modification, the speaker identification unit 14 identifies the speaker of the utterance data based on the score representing the weighted similarity output by the similarity calculation unit 135B. More specifically, the speaker identification unit 14 identifies the speaker of the entire utterance data using the overall score, which is the arithmetic average of the scores for each of the multiple utterance data output by the reliability assignment unit 138B. Here, the speaker identification unit 14 identifies the speaker of the entire utterance using the overall scores that are equal to or greater than a threshold value. Then, the speaker identification unit 14 outputs the identified speaker of the entire utterance as an identification result. In this way, the speaker identification unit 14 can accurately identify the speaker of the entire utterance data corresponding to the overall score by using only the highly reliable overall scores.
[0106] [Processing example of speaker identification device 11B] Next, an example of the processing of the speaker identification device 11B configured as above will be described with reference to FIG.
[0107] FIG. 11 is a diagram showing an example of processing by the speaker identification device 11B according to Modification 2 of this embodiment. The top row of FIG. 11 shows all speech data acquired by the speaker identification device 11B. As described above, the all speech data is a voice signal converted from the voice of one speaker over a predetermined period, and is composed of speech data divided into each identification unit. In the example shown in FIG. 11, the identification unit is, for example, an interval of 3 to 4 seconds, and the all speech data is a voice signal of a voice lasting 12 to 16 seconds, and is divided into voice signals of four identification units. The all speech data divided into each identification unit corresponds to the above-mentioned speech data.
[0108] The second row of Fig. 11 shows the pre-weighted scores and estimation results for each of the multiple utterance data. The pre-weighted scores represent the similarity between each of the multiple utterance data calculated by the speaker identification device 11B. The estimation results represent the emotions included in the voice of the utterance indicated by the utterance data, estimated by the speaker identification device 11B for each of the multiple utterance data constituting the entire utterance data. In the example shown in Fig. 11, (score, emotion) is shown as (50, calm), (50, angry), (50, whisper), and (50, angry) for each identification unit of the entire utterance data (each utterance data).
[0109] 11 shows a score weighted based on the estimation result. This score is a weighted similarity based on the estimation result for each of the plurality of utterance data, and represents the similarity for each of the plurality of utterance data. In the example shown in FIG. 11, the largest weight is assigned when the emotion indicated by the estimation result is “neutral,” and the weights are 75, 25, 5, and 25 for each identification unit (each utterance data) of the entire utterance data. Note that the largest weight is assigned when the emotion indicated by the estimation result is “neutral.” This is because the speaker identification device 11B calculates the similarity for each of the plurality of utterance data using the second speaker feature of the registered utterance containing the “neutral” emotion of the registered speaker. In other words, the greater the match with the emotion that may be included in the registered utterance used to obtain the second speaker feature used by the speaker identification device 11B when calculating the similarity, the higher the reliability of the calculated similarity is considered to be, and the larger the weight is assigned.
[0110] The fourth row of Fig. 11 shows the overall score. The overall score is the score for the entire utterance data, and is the arithmetic average of the scores for each of the multiple utterance data, as described above. In the example shown in Fig. 11, the overall score is calculated as 32.5.
[0111] [Effects, etc.] As described above, in the speaker identification device 11B according to this modification, the speaker identification processing unit 13B outputs a score obtained by assigning a weight based on the estimation result of the emotion of the evaluation utterance to the similarity calculated between the evaluation utterance and the registered utterance. Note that the speaker identification processing unit 13B assigns a larger weight to the degree of similarity calculated, since the reliability of the calculated similarity is higher as the emotion included in the evaluation utterance indicated by the estimation result matches the emotion included in the registered utterance.
[0112] As a result, by using a highly reliable score, it is possible to identify the speaker of the evaluation utterance when the emotion included in the registration utterance and the emotion included in the evaluation utterance are close (similar).
[0113] Therefore, the speaker identification device 11B according to this modification can improve the accuracy of speaker identification even if the speaker's emotion is included in the utterance to be identified.
[0114] The reliability of the speaker identification result may be confirmed by checking the reliability of the score.
[0115] (Variation 3) In the second modification, the speaker identification device 11B outputs a score obtained by assigning a weight based on the estimation result of the emotion included in the evaluation utterance as reliability to the calculated similarity. In the third modification, the speaker identification device 11C assigns a reliability (specifically, additional information indicating the reliability) based on the estimation result of the emotion included in the evaluation utterance to the calculated similarity and outputs the score. The following describes the speaker identification device 11C according to the third modification, focusing on the differences from the speaker identification device 11B described in the second modification.
[0116] [8. Speaker Identification Device 11C] Fig. 12 is a block diagram showing an example of a detailed configuration of a speaker identification device 11C according to Modification 3 of the present embodiment. Note that the same elements as those in Figs. 4, 9, 10, etc. are given the same reference numerals, and detailed description thereof will be omitted.
[0117] Similar to the speaker identification device 11B, the speaker identification device 11C performs processing to identify the speaker of utterance data indicating the voice of the utterance to be identified. More specifically, the speaker identification device 11C calculates a score representing the similarity between a first speaker feature of the evaluation utterance and a second speaker feature of the registered utterance for each registered speaker. Then, the speaker identification device 11B outputs the score obtained by adding a reliability (or additional information representing the reliability) to the calculated similarity as the identification result.
[0118] As shown in Fig. 12, a speaker identification device 11C according to this modification differs from the speaker identification device 11B shown in Fig. 10 in the configuration of a speaker identification processing unit 13C. More specifically, the speaker identification device 11C according to this modification differs from the speaker identification device 11B shown in Fig. 10 in the configuration in that it does not have a reliability assigning unit 138B but has a reliability assigning unit 138C.
[0119] [8.1 Trust Assignment Unit 138C] The confidence assigning unit 138C assigns a confidence level to the similarity calculated by the similarity calculation unit 135B according to the emotion indicated by the estimation result of the emotion estimator 12, and outputs the result as a score. Here, when one emotion matches the emotion indicated by the estimation result, the confidence assigning unit 138C assigns the highest confidence level to the calculated similarity.
[0120] [9. Speaker Identification Unit 14] The speaker identification unit 14 uses the score output by the speaker identification device 11C to identify the speaker of the speech data.
[0121] In this modification, the speaker identification unit 14 identifies the speaker of the utterance data based on the score indicating the similarity to which a reliability has been assigned, which has been output by the similarity calculation unit 135B. For example, the speaker identification unit 14 identifies the speaker of the utterance data using a score to which a reliability equal to or greater than a threshold has been assigned. Then, the speaker identification unit 14 outputs the speaker of the identified utterance as an identification result. In this way, the speaker identification unit 14 can accurately identify the speaker of the utterance data corresponding to the score by using only highly reliable scores.
[0122] [Effects, etc.] As described above, in the speaker identification device 11C according to this modification, the speaker identification processing unit 13C outputs a score obtained by adding additional information indicating reliability based on the estimation result of the emotion of the evaluation utterance to the similarity calculated between the evaluation utterance and the registered utterance. For example, the speaker identification processing unit 13C adds additional information such that the reliability of the calculated similarity increases as the emotion included in the evaluation utterance indicated by the estimation result matches the emotion included in the registered utterance.
[0123] As a result, by using a highly reliable score, it is possible to identify the speaker of the evaluation utterance when the emotion included in the registration utterance and the emotion included in the evaluation utterance are close (similar).
[0124] Therefore, the speaker identification device 11C according to this modification can improve the accuracy of speaker identification even if the speech to be identified contains the emotion of the speaker.
[0125] The reliability of the speaker identification result may be confirmed by checking the reliability of the score.
[0126] (Possibilities for other embodiments) Although the speaker identification devices according to the embodiments and modifications have been described above, the present disclosure is not limited to these embodiments.
[0127] For example, each processing unit included in the speaker identification device according to the above-described embodiment and modifications is typically realized as an LSI, which is an integrated circuit. These units may be individually implemented as single chips, or some or all of them may be integrated into a single chip.
[0128] Furthermore, the integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connections and settings of circuit cells within LSIs to be reconfigured, may also be used.
[0129] The present disclosure may also be realized as a speaker identification method executed by a speaker identification device.
[0130] In each of the above embodiments, each component may be configured with dedicated hardware, or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.
[0131] The division of functional blocks in the block diagram is an example, and multiple functional blocks may be realized as a single functional block, one functional block may be divided into multiple blocks, or some functions may be moved to another functional block.Furthermore, the functions of multiple functional blocks having similar functions may be processed in parallel or time-shared by a single piece of hardware or software.
[0132] The order in which the steps in the flowchart are executed is merely an example for specifically explaining the present disclosure, and other orders may be used. Some of the steps may be executed simultaneously (in parallel) with other steps.
[0133] The speaker identification device according to one or more aspects has been described above based on the embodiments and modifications, etc., but the present disclosure is not limited to these embodiments and modifications, etc. As long as it does not deviate from the spirit of the present disclosure, various modifications that a person skilled in the art may make to these embodiments and modifications, etc., or configurations constructed by combining components of different embodiments and modifications, etc., may also be included within the scope of one or more aspects. [Industrial Applicability]
[0134] The present disclosure can be used in a speaker identification device, a speaker identification method, and a program, and can be used in, for example, a meeting minutes system, a communication visualization system, or the like, for a speaker identification device, a speaker identification method, and a program that identify the speaker of spontaneous speech that includes emotions. [Explanation of symbols]
[0135] 1. Speaker Identification System 10 Pretreatment section 11, 11A, 11B, 11C Speaker identification device 12 Emotion Estimator 13, 13A, 13B, 13C Speaker identification processing unit 14 Speaker Identification Unit 101 Voice Acquisition Unit 102 Acoustic feature calculation unit 121, 1331 Frame connection processing unit 122, 1332b DNN 123, 134A, 134b, 136A, 1333b Storage section 131 Classifier selection unit 132, 132a, 132b Speaker classifier 133A, 133b Speaker feature extraction unit 135A, 135B, 135b Similarity calculation section 137A Correction section 138B, 138C Reliability Granting Section
Claims
1. A speaker identification device for identifying a speaker of utterance data indicating a speech of an utterance to be identified, comprising: an emotion estimator that estimates an emotion included in the speech represented by the speech data from acoustic features calculated from the speech data by using a trained DNN (Deep Neural Network); a speaker identification processing unit that calculates a similarity between a second speaker feature pre-registered for each registered user that corresponds to an emotion indicated by an estimation result of the emotion estimator and a first speaker feature of an evaluation utterance extracted from the acoustic feature calculated from the utterance data, and outputs a score for identifying the speaker of the utterance data based on the similarity. Speaker identification device.
2. A speaker identification device for identifying a speaker of utterance data indicating a speech of an utterance to be identified, comprising: an emotion estimator that estimates an emotion included in the speech represented by the speech data from acoustic features calculated from the speech data by using a trained DNN (Deep Neural Network); a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data, The speaker identification processing unit a plurality of speaker classifiers each including: a speaker feature extraction unit that, when the acoustic features are input, extracts first speaker features that can identify a speaker of the speech represented by the speech data from the input acoustic features; and a similarity calculation unit that calculates a similarity between the first speaker features extracted by the speaker feature extraction unit and second speaker features that are stored in a storage unit and can identify each speech including one emotion of a registered speaker that is a classification target; a classifier selection unit that selects one of the plurality of speaker classifiers, the speaker classifier having stored in the storage unit a second speaker feature that can identify each speech containing one emotion of the registered speaker corresponding to the emotion indicated by the estimation result, the speaker classifier selected by the classifier selection unit receives the acoustic features calculated from the speech data, calculates the similarity, and outputs the similarity as the score. Speaker identification device.
3. A speaker identification device for identifying a speaker of utterance data indicating a speech of an utterance to be identified, comprising: an emotion estimator that estimates an emotion included in the speech represented by the speech data from acoustic features calculated from the speech data by using a trained DNN (Deep Neural Network); a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data, The speaker identification processing unit a speaker feature extraction unit that extracts, from the acoustic feature, a first speaker feature that can identify a speaker of the voice of the utterance indicated by the utterance data; a correction unit that corrects second speaker features stored in the storage unit, the second speaker features being capable of identifying each speech containing one emotion of a registered speaker to be identified, into third speaker features that can identify each speech containing the one emotion corresponding to the emotion indicated by the estimation result; a similarity calculation unit that calculates a similarity between the extracted first speaker feature and the third speaker feature corrected by the correction unit, and outputs the calculated similarity as the score. Speaker identification device.
4. A speaker identification device for identifying a speaker of utterance data indicating a speech of an utterance to be identified, comprising: an emotion estimator that estimates an emotion included in the speech represented by the speech data from acoustic features calculated from the speech data by using a trained DNN (Deep Neural Network); a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data, The speaker identification processing unit a speaker feature extraction unit that extracts, from the acoustic feature, a first speaker feature that can identify a speaker of the voice of the utterance indicated by the utterance data; a similarity calculation unit that calculates a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each speech including one emotion of a registered speaker that is an identification target; a reliability assigning unit that assigns a weight to the calculated similarity according to the emotion indicated by the estimation result and outputs the weight as the score, the reliability assigning unit assigns the largest weight to the calculated similarity when the one emotion and the emotion indicated by the estimation result match. Speaker identification device.
5. the acoustic feature is calculated from each of a plurality of pieces of speech data acquired by dividing entire speech data representing speech of one speaker over a predetermined period into identification units in a time series by a preprocessing unit; the reliability assigning unit assigns a weight to the similarity for each of the plurality of utterance data calculated by the similarity calculation unit according to an emotion indicated by the estimation result for each of the plurality of utterance data estimated by the emotion estimator, and outputs the weight as the score.
5. The speaker identification device according to claim 4.
6. The speaker identification device further comprises: a speaker identification unit that identifies a speaker of the entire utterance data using an overall score that is an arithmetic average of the scores for each of the plurality of utterance data output by the reliability assignment unit, the speaker identification unit identifies a speaker of the entire speech data by using an overall score equal to or greater than a threshold among the overall scores.
6. The speaker identification device according to claim 5.
7. A speaker identification device for identifying a speaker of utterance data indicating a speech of an utterance to be identified, comprising: an emotion estimator that estimates an emotion included in the speech represented by the speech data from acoustic features calculated from the speech data by using a trained DNN (Deep Neural Network); a speaker identification processing unit that uses an estimation result of the emotion estimator to output a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data, The speaker identification processing unit a speaker feature extraction unit that extracts, from the acoustic feature, a first speaker feature that can identify a speaker of the voice of the utterance indicated by the utterance data; a similarity calculation unit that calculates a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each speech including one emotion of a registered speaker that is an identification target; a reliability assigning unit that assigns a reliability to the calculated similarity according to the emotion indicated by the estimation result and outputs the result as the score, Speaker identification device.
8. The speaker identification device further comprises: a speaker identification unit that identifies a speaker of the speech data using the score where the reliability is equal to or greater than a threshold value; 8. The speaker identification device according to claim 7.
9. the speaker feature extraction unit extracts the first speaker feature from the acoustic feature using a trained DNN; The speaker identification device according to any one of claims 2 to 8.
10. A speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, comprising: an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of calculating a similarity between a second speaker feature pre-registered for each registered user that corresponds to the emotion indicated by the estimation result in the emotion estimation step and a first speaker feature of an evaluation utterance extracted from the acoustic feature calculated from the utterance data, and outputting a score for identifying the speaker of the utterance data based on the similarity. Speaker identification methods.
11. A speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, comprising: an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, selecting one of the plurality of speaker classifiers, the speaker classifier having stored in a storage unit a second speaker feature that can identify each speech containing one emotion of the registered speaker to be classified, according to the emotion indicated by the estimation result; Each of the plurality of speaker classifiers includes a speaker feature extraction unit that, when the acoustic features are input, extracts, from the input acoustic features, first speaker features that can identify a speaker of the speech of the utterance indicated by the speech data, and a similarity calculation unit that calculates a similarity between the first speaker features extracted by the speaker feature extraction unit and second speaker features that are stored in a storage unit and can identify each speech containing one emotion of a registered speaker that is to be classified, The selected speaker classifier receives the acoustic features calculated from the speech data, calculates the similarity, and outputs it as the score. Speaker identification methods.
12. A speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, comprising: an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; correcting second speaker features stored in a storage unit, the second speaker features being capable of identifying each speech containing one emotion of a registered speaker to be identified, to third speaker features being capable of identifying each speech containing the one emotion corresponding to the emotion indicated by the estimation result; calculating a similarity between the extracted first speaker feature and the corrected third speaker feature, and outputting the calculated similarity as the score; Speaker identification methods.
13. A speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, comprising: an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; calculating a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each of the speeches including one emotion of the registered speaker to be identified; a weighting factor according to the emotion indicated by the estimation result is assigned to the calculated similarity, and the weighting factor is output as the score; The highest weight is assigned to the similarity calculated when the one emotion and the emotion indicated by the estimation result match. Speaker identification methods.
14. A speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, comprising: an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; calculating a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each of the speeches including one emotion of the registered speaker to be identified; a reliability corresponding to the emotion indicated by the estimation result is assigned to the calculated similarity, and the resulting reliability is output as the score; Speaker identification methods.
15. A program for causing a computer to execute a speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of calculating a similarity between a second speaker feature pre-registered for each registered user that corresponds to the emotion indicated by the estimation result in the emotion estimation step and a first speaker feature of an evaluation utterance extracted from the acoustic feature calculated from the utterance data, and outputting a score for identifying the speaker of the utterance data based on the similarity; program.
16. A program for causing a computer to execute a speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, selecting one of the plurality of speaker classifiers, the speaker classifier having stored in a storage unit a second speaker feature that can identify each speech containing one emotion of the registered speaker to be classified, according to the emotion indicated by the estimation result; Each of the plurality of speaker classifiers includes a speaker feature extraction unit that, when the acoustic features are input, extracts, from the input acoustic features, first speaker features that can identify a speaker of the speech of the utterance indicated by the speech data, and a similarity calculation unit that calculates a similarity between the first speaker features extracted by the speaker feature extraction unit and second speaker features that are stored in a storage unit and can identify each speech containing one emotion of a registered speaker that is to be classified, The selected speaker classifier receives the acoustic features calculated from the speech data, calculates the similarity, and outputs it as the score. causing a computer to perform a speaker identification method; program.
17. A program for causing a computer to execute a speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; correcting second speaker features stored in a storage unit, the second speaker features being capable of identifying each speech containing one emotion of a registered speaker to be identified, to third speaker features being capable of identifying each speech containing the one emotion corresponding to the emotion indicated by the estimation result; calculating a similarity between the extracted first speaker feature and the corrected third speaker feature, and outputting the calculated similarity as the score; causing a computer to perform a speaker identification method; program.
18. A program for causing a computer to execute a speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; calculating a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each of the speeches including one emotion of the registered speaker to be identified; a weighting factor according to the emotion indicated by the estimation result is assigned to the calculated similarity, and the weighting factor is output as the score; The highest weight is assigned to the similarity calculated when the one emotion and the emotion indicated by the estimation result match. causing a computer to perform a speaker identification method; program.
19. A program for causing a computer to execute a speaker identification method for identifying a speaker of utterance data indicating the voice of an utterance to be identified, an emotion estimation step of estimating an emotion included in the speech of the utterance indicated by the utterance data from acoustic features calculated from the utterance data using a trained DNN; a speaker identification processing step of outputting a score for identifying a speaker of the utterance data from the acoustic feature calculated from the utterance data using a result of the estimation in the emotion estimation step, In the speaker identification processing step, extracting, from the acoustic features, first speaker features that can identify a speaker of the voice of the utterance indicated by the utterance data; calculating a similarity between the extracted first speaker feature and a second speaker feature stored in a storage unit, the second speaker feature being capable of identifying each of the speeches including one emotion of the registered speaker to be identified; a reliability corresponding to the emotion indicated by the estimation result is assigned to the calculated similarity, and the resulting reliability is output as the score; causing a computer to perform a speaker identification method; program.
Citation Information
Patent Citations
Emotion recognition method and device based on short video voice
CN110473571A
Registered utterance division device, speaker likelihood evaluation device, speaker identification device, registered utterance division method, speaker likelihood evaluation method, and program
JP2017187642A
Voice processing program, voice processing method and voice processor
JP2020126125A