Speaker identification apparatus, speaker identification method, and recording medium
By using deep neural networks to infer emotion and calculate acoustic features, the problem of emotional speech affecting recognition accuracy is solved, and efficient speaker recognition is achieved even when emotion is present.
Patent Information
- Application Number
- CN202180013727.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-31
- Filing Date
- 2021-02-05
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2041-02-05
AI Technical Summary
Existing speaker recognition technologies suffer from decreased accuracy when faced with emotional speech, especially when the emotions in the registered speech and the evaluation speech differ.
A deep neural network (DNN) is used for sentiment inference, combined with a speaker recognition processing unit. By calculating acoustic features and sentiment inference results, the recognition accuracy is improved.
Even when the speech contains emotion, it can effectively improve the accuracy of speaker recognition, and is suitable for free speech recognition in meeting minutes and communication visualization systems.
Smart Images

Figure CN115104152B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to a speaker identification device, a speaker identification method, and a recording medium. BACKGROUND
[0002] Speaker identification technology is a technology of estimating a registered speech which is a speech of each speaker as a registration object, based on similarity of a feature quantity calculated from the registered speech and a feature quantity calculated from an evaluation speech which is a speech of an unknown speaker as an identification object, and identifying a speaker of the evaluation speech (for example, Patent Literature 1).
[0003] For example, as the speaker identification technology, Patent Literature 1 discloses a technology of identifying a speaker of an evaluation speech by using similarity of a speaker feature vector in a registered speech of each registered speaker and a speaker feature vector in the evaluation speech.
[0004] (Patent Literature)
[0005] (Patent Literature)
[0006] Patent Literature 1: Japanese Patent Application Publication No. 2017-187642
[0007] However, in a case where emotional speech such as laughter or an angry shout is set as the evaluation speech, the recognition accuracy is affected. Specifically, if the emotion included in the registered speech is different from the emotion included in the evaluation speech, a voice intonation change occurs as the emotion included in the evaluation speech is different, which can cause a decrease in the recognition accuracy of the speaker.
[0008] That is, in the existing speaker identification technology disclosed in Patent Literature 1, the similarity of the speaker feature vectors in the registered speech and the evaluation speech is calculated without considering the emotion included in the evaluation speech, and the speaker of the evaluation speech is identified. Therefore, with the current speaker identification technology, the accuracy of identifying the speaker of the evaluation speech is sometimes insufficient. SUMMARY
[0009] In view of the above problem, an object of the present disclosure is to provide a speaker identification device, a speaker identification method, and a recording medium capable of improving the recognition accuracy of a speaker even if the emotion of the speaker is included in an evaluation speech which is an identification object.
[0010] One aspect of the present disclosure relates to a speaker identification device that identifies a speaker corresponding to speech data that shows a speech sound of an identification target, the speaker identification device including an emotion estimator that estimates an emotion included in the speech sound shown by the speech data based on an acoustic feature quantity calculated from the speech data, using a DNN (Deep Neural Network) that has been learned, and a speaker identification processing section that outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantity calculated from the speech data, using a result of the estimation by the emotion estimator.
[0011] In addition, these general or specific aspects can also be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a CD-ROM, and any combination of the system, the method, the integrated circuit, the computer program, and the recording medium.
[0012] With the speaker identification device and the like of the present disclosure, even if an emotion of a speaker is included in speech of an identification target, it is possible to improve the accuracy of identification of the speaker. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is a block diagram showing one example of a configuration of a speaker identification system according to an embodiment.
[0014] Figure 2 is a block diagram showing another example of a configuration of a speaker identification system according to an embodiment.
[0015] Figure 3 is a block diagram showing one example of a detailed configuration of a pre-processing section according to an embodiment.
[0016] Figure 4 is a block diagram showing one example of a detailed configuration of a speaker identification device according to an embodiment.
[0017] Figure 5 shows one example of a configuration of an emotion estimator according to an embodiment.
[0018] Figure 6 shows one example of a configuration of a speaker identifier according to an embodiment.
[0019] Figure 7 shows one example of a configuration of a speaker feature quantity extractor included in a speaker identifier according to an embodiment.
[0020] Figure 8 is a flowchart showing an outline of the operation of a speaker identification device according to an embodiment.Figure 9 is a block diagram showing one example of a specific configuration of a speaker recognition device according to Embodiment 1.
[0021] Figure 10 is a block diagram showing one example of a specific configuration of a speaker recognition device according to Embodiment 2.
[0022] Figure 11 One processing example of the speaker recognition device according to Embodiment 2 is shown.
[0023] Figure 12 is a block diagram showing one example of a specific configuration of a speaker recognition device according to Embodiment 3. DETAILED DESCRIPTION
[0024] (SUMMARY OF THE DISCLOSURE)
[0025] A summary of one aspect of the present disclosure is described below.
[0026] The speaker recognition device according to one aspect of the present disclosure identifies a speaker corresponding to speech data showing a speech sound of a recognition target, and includes an emotion estimator that estimates an emotion included in the speech sound shown by the speech data, based on an acoustic feature quantity calculated from the speech data, using a DNN (Deep Neural Network) that has been subjected to learning, and a speaker recognition processing section that outputs a score for identifying the speaker corresponding to the speech data, based on the acoustic feature quantity calculated from the speech data, using a result of the estimation by the emotion estimator.
[0027] According to this aspect, even if an emotion of a speaker is included in a speech of a recognition target, it is possible to improve the recognition accuracy of the speaker.
[0028] Also, for example, the speaker recognition processing section can include a plurality of speaker recognizers each having a speaker feature amount extraction section that extracts a first speaker feature amount capable of determining a speaker of the speech sound shown by the speech data from the input acoustic feature amount, and a similarity calculation section that calculates a similarity of the first speaker feature amount extracted by the speaker feature amount extraction section and a second speaker feature amount stored in a storage section, and the second speaker feature amount is a feature amount capable of determining each of sounds containing one emotion of a registered speaker as an object of recognition, and an recognizer selection section that selects one speaker recognizer from the plurality of speaker recognizers, and the selected one speaker recognizer is a speaker recognizer in which the second speaker feature amount capable of determining each of sounds corresponding to the emotion shown by the estimation result, containing one emotion of the registered speaker, is stored in the storage section, and the speaker recognizer selected by the recognizer selection section calculates the similarity by inputting the acoustic feature amount calculated from the speech data and outputs the similarity as the score.
[0029] Also, for example, the speaker recognition processing section can include a speaker feature amount extraction section that extracts a first speaker feature amount capable of determining a speaker of the speech sound shown by the speech data from the input acoustic feature amount, a modification section that modifies a second speaker feature amount stored in a storage section to a third speaker feature amount, the second speaker feature amount being capable of determining each of sounds containing one emotion of a registered speaker as an object of recognition, the third speaker feature amount being capable of determining each of sounds containing one emotion corresponding to the emotion shown by the estimation result, and a similarity calculation section that calculates a similarity of the extracted first speaker feature amount and the third speaker feature amount modified by the modification section, and outputs the calculated similarity as the score.
[0030] Also, for example, the speaker recognition processing section can include a speaker feature quantity extraction section that extracts a first speaker feature quantity from the acoustic feature quantity, the first speaker feature quantity being capable of determining a speaker of the speech sound shown by the speech data, a similarity calculation section that calculates a similarity of the extracted first speaker feature quantity and a second speaker feature quantity stored in a storage section, the second speaker feature quantity being a feature quantity capable of determining each of sounds containing a kind of emotion of a registered speaker who is an object of recognition, and a reliability assignment section that assigns a weight corresponding to an emotion shown by the estimation result to the calculated similarity, and outputs the weight as the score. The reliability assignment section can assign a maximum weight to the calculated similarity when the kind of emotion coincides with the emotion shown by the estimation result.
[0031] Also, for example, the acoustic feature quantity can be calculated by a pre-processing section dividing, in time series and by recognition units, all the speech data showing speech sounds of one speaker in a prescribed period, thereby obtaining a plurality of speech data, and the acoustic feature quantity can be calculated from each of the obtained plurality of speech data. The reliability assignment section can assign a weight to the similarity calculated by the similarity calculation section for each of the plurality of speech data, the weight being a weight corresponding to an emotion shown by the estimation result estimated by the emotion estimator for each of the plurality of speech data.
[0032] Also, for example, the speaker recognition apparatus can further include a speaker recognition section that recognizes a speaker corresponding to the all speech data using a total score, the total score being a score obtained by arithmetically averaging the scores output by the reliability assignment section for each of the plurality of speech data, and the speaker recognition section can recognize the speaker corresponding to the all speech data using a total score of the total scores that is equal to or higher than a threshold value.
[0033] Also, for example, the speaker recognition processing section can include a speaker feature quantity extraction section that extracts a first speaker feature quantity from the acoustic feature quantity, the first speaker feature quantity being capable of determining a speaker of the speech sound shown by the speech data, a similarity calculation section that calculates a similarity of the extracted first speaker feature quantity and a second speaker feature quantity stored in a storage section, the second speaker feature quantity being a feature quantity capable of determining each of sounds containing a kind of emotion of a registered speaker who is an object of recognition, and a reliability assignment section that assigns a weight corresponding to an emotion shown by the estimation result to the calculated similarity, and outputs the weight as the score. The reliability assignment section can assign a maximum weight to the calculated similarity when the kind of emotion coincides with the emotion shown by the estimation result.
[0034] Further, for example, the speaker identification apparatus can further include a speaker identification unit that identifies a speaker corresponding to the speech data using the score that is equal to or higher than the threshold.
[0035] Further, for example, the speaker feature quantity extraction unit can extract the first speaker feature quantity from the acoustic feature quantity using the DNN that has been learned.
[0036] A speaker identification method according to one aspect of the present disclosure identifies a speaker corresponding to speech data that shows speech sound that is an object of identification, the speaker identification method including: an emotion estimation step of estimating an emotion included in the speech sound shown by the speech data using a DNN that has been learned, based on an acoustic feature quantity calculated from the speech data; and a speaker identification processing step of outputting a score for identifying the speaker corresponding to the speech data, based on the acoustic feature quantity calculated from the speech data, using an estimation result of the emotion estimator.
[0037] Further, a recording medium according to one aspect of the present disclosure is a non-transitory recording medium that is readable by a computer, in which a program is recorded, the program causing the computer to execute a speaker identification method that identifies a speaker corresponding to speech data that shows speech sound that is an object of identification, the speaker identification method including: an emotion estimation step of estimating an emotion included in the speech sound shown by the speech data using a DNN that has been learned, based on an acoustic feature quantity calculated from the speech data; and a speaker identification processing step of outputting a score for identifying the speaker corresponding to the speech data, based on the acoustic feature quantity calculated from the speech data, using an estimation result of the emotion estimator.
[0038] Further, these general or specific aspects can be implemented by a system, a method, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, and any combination of the system, the method, the integrated circuit, the computer program, and the recording medium.
[0039] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. The embodiments to be described below are merely one specific example of the present disclosure. The numerical values, shapes, constituent elements, steps, and order of steps and the like shown in the following embodiments are merely one example, and the gist of the present disclosure is not limited thereto. Also, for the constituent elements of the following embodiments that are not described in the constituent elements of the independent technical solution showing the most general concept, the constituent elements will be described as arbitrary constituent elements. Also, in all of the embodiments, the contents in each of the embodiments can be combined.
[0040] (Embodiment)
[0041] A speaker recognition device and the like according to the present embodiment will be described below with reference to the accompanying drawings.
[0042] [Speaker recognition system 1]
[0043] Figure 1 is a block diagram showing one example of the configuration of the speaker recognition system 1 according to the present embodiment. Figure 2 is a block diagram showing another example of the configuration of the speaker recognition system 1 according to the present embodiment.
[0044] The speaker recognition system 1 according to the present embodiment is used to identify a speaker to which the speech data showing the speech of the speaker containing the emotion of the speaker corresponds, and here, the speech of the speaker containing the emotion of the speaker is the speech to be recognized. As shown in Figure 1 , the speaker recognition system 1 is provided with a preprocessing section 10 and a speaker recognition device 11. Also, as shown in Figure 2 , the speaker recognition system 1 can be provided with a speaker recognition section 14, but such a configuration is not necessary. The constituent elements will be described below.
[0045] [1. Preprocessing section 10]
[0046] Figure 3 is a block diagram showing one example of the detailed configuration of the preprocessing section 10 according to the present embodiment.
[0047] The preprocessing section 10 obtains the speech data showing the speech to be recognized, and outputs the acoustic feature quantity calculated from the obtained speech data to the speaker recognition device 11. As shown in Figure 3 , the preprocessing section 10 according to the present embodiment is provided with a sound acquisition section 101 and an acoustic feature quantity calculation section 102.
[0048] [1.1 Sound acquisition section 101]
[0049] The sound acquisition unit 101 is configured by, for example, a microphone, and acquires a speech sound of a speaker. The sound acquisition unit 101 converts the acquired speech sound into a sound signal, detects a speech interval that is an interval in which a speech has been made, and outputs speech data that shows the speech sound obtained by dividing the speech interval to the acoustic feature quantity calculation unit 102.
[0050] In addition, the sound acquisition unit 101 can divide all the speech data that shows the speech sound of one speaker in a prescribed period in a time series and by a recognition unit, thereby obtaining a plurality of speech data, and output the plurality of speech data to the acoustic feature quantity calculation unit 102. The recognition unit can be, for example, a time length of 3 to 4 seconds, or the above-described speech interval.
[0051] [1.2 Acoustic Feature Quantity Calculation Unit 102]
[0052] The acoustic feature quantity calculation unit 102 calculates an acoustic feature quantity related to a speech sound, based on the sound signal, that is, the speech data of the speech interval output from the sound acquisition unit 101. The acoustic feature quantity calculation unit 102 in the present embodiment calculates MFCC (Mel Frequency Cepstral Coefficient) that is a feature quantity of a speech sound, based on the speech data output from the sound acquisition unit 101, as an acoustic feature quantity. The MFCC is a feature quantity that represents a vocal tract feature of a speaker, and is generally used for voice recognition. More specifically, the MFCC is an acoustic feature quantity obtained by analyzing a sound spectrum based on a human auditory feature. In addition, as the acoustic feature quantity, it is not limited to the case where the acoustic feature quantity calculation unit 102 calculates the MFCC from the speech data, but it can be calculated by filtering the sound signal with a mel filter bank, and it can be calculated as a spectrum of the sound signal as the acoustic feature quantity.
[0053] [2. Speaker Recognition Device 11]
[0054] The speaker recognition device 11 can be realized by, for example, a computer that has a processor (microprocessor), a memory, a communication interface, and the like. The speaker recognition device 11 can be operated in a server, or a part of the speaker recognition device 11 can be operated in a cloud server. The speaker recognition device 11 performs processing for identifying a speaker to which speech data corresponds, the speech data showing a speech sound that is an evaluation target of recognition. More specifically, the speaker recognition device 11 outputs, as an identification result, a score that shows a similarity between a first speaker feature quantity that shows an evaluation speech of a speaker and a second speaker feature quantity that shows a registered speech of each registered speaker. The evaluation speech, that is, the speech that is an evaluation target of recognition in the present embodiment includes a speaker's emotion.
[0055] Figure 4 is a block diagram showing one example of a specific configuration of the speaker recognition device 11 according to the present embodiment.
[0056] As shown in Figure 1 and Figure 4 , the speaker recognition device 11 includes an emotion estimator 12 and a speaker recognition processing section 13.
[0057] [2.1 Emotion Estimator 12]
[0058] The emotion estimator 12 estimates an emotion included in a speech sound shown by speech data, based on an acoustic feature quantity calculated from the speech data, using a DNN (Deep Neural Network) that has been subjected to learning. Note that, as the DNN, a CNN (Convolution Neural Networks), a fully connected NN (Neural Network), or a TDNN (Time Delay Neural Network) can be used, for example.
[0059] Here, one example of a configuration of the emotion estimator 12 will be described with reference to Figure 5 .
[0060] Figure 5 One example of a configuration of the emotion estimator 12 according to the present embodiment will be described.
[0061] As shown in Figure 5 , the emotion estimator 12 includes a frame connection processing section 121 and a DNN 122.
[0062] [2.1.1 Frame Connection Processing Section 121]
[0063] The frame connection processing section 121 connects a plurality of frames of MFCCs, which are acoustic feature quantities output from the pre-processing section 10, and outputs the same to an input layer of the DNN 122. The MFCCs are composed of a plurality of frames each having x (x is a positive integer)th power of a feature quantity. In the example shown in Figure 5 , the frame connection processing section 121 connects MFCC parameters composed of 24th power of a feature quantity per frame into 50 frames, thereby generating a 1200th power vector, which is output to the input layer of the DNN 122.
[0064] [2.1.2 DNN 122]
[0065] When a plurality of frames of the connected MFCCs are input, the DNN 122 outputs a label of an emotion with the highest probability as an estimation result of the emotion estimator 12. In the example shown in Figure 5In the example, the DNN 122 is a neural network constituted by an input layer, a plurality of intermediate layers, and an output layer, and is a neural network that has been learned using teacher data stored in the storage section 123, the teacher data being teacher voice data containing the emotion of the estimation target. The input layer is constituted by, for example, 1200 nodes, and a 1200-dimensional vector generated by connecting MFCC parameters constituted by 24-dimensional feature amounts for each frame into 50 frames is input to the input layer. The output layer is constituted by, for example, nodes outputting emotion labels such as calm, anger, laughter, and sadness, and outputs the emotion label with the highest probability. In addition, the plurality of intermediate layers is constituted by, for example, 2 to 3 layers.
[0066] [2.2 Speaker identification processing section 13]
[0067] The speaker identification processing section 13 outputs a score for identifying the speaker corresponding to the speech data, using the estimation result of the emotion estimator 12, based on the acoustic feature amount calculated from the speech data.
[0068] As shown in Figure 4 , the speaker identification processing section 13 in the present embodiment is provided with an identifier selection section 131 and a plurality of speaker identifiers 132.
[0069] [2.2.1 Plurality of speaker identifiers 132]
[0070] Each of the plurality of speaker identifiers 132 is a speaker identifier 132k (k is a natural number) corresponding to one emotion. One emotion refers to, for example, one of calm, anger, laughter, and sadness. In Figure 4 the example, the plurality of speaker identifiers 132 is constituted by a speaker identifier 132a, a speaker identifier 132b, and the like. For example, the speaker identifier 132a corresponds to "calm" as one emotion, and the speaker identifier 132b corresponds to "laughter" as one emotion. In addition, one speaker identifier among the speaker identifier 132a, the speaker identifier 132b, and the like is referred to as the speaker identifier 132k.
[0071] When the speaker identifier 132k selected by the identifier selection section 131 among the plurality of speaker identifiers 132 is input with the acoustic feature amount calculated from the speech data, the similarity is calculated and output as a score. In addition, there is a case where none of the plurality of speaker identifiers 132 is selected by the identifier selection section 131, and for the case where none of the speaker identifiers 132 is selected by the identifier selection section 131, "no selection" is indicated in Figure 4 .
[0072] Here, as one example of the speaker identifier 132k, the speaker identifier 132b corresponding to "laughter" will be described with reference to Figure 6 .
[0073] Figure 6 An example of the configuration of the speaker recognizer 132b according to the present embodiment is shown. Figure 7 An example of the configuration of the speaker feature amount extraction section 133b included in the speaker recognizer 132b according to the present embodiment is shown.
[0074] As shown in Figure 6 , the speaker recognizer 132b includes a speaker feature amount extraction section 133b, a storage section 134b, and a similarity calculation section 135b.
[0075] [2.2.1.1 Speaker feature amount extraction section 133b]
[0076] The speaker feature amount extraction section 133b extracts, from the acoustic feature amount calculated from the speech data, a first speaker feature amount that enables determination of the speaker of the speech sound shown in the speech data, in a case where the acoustic feature amount is input. More specifically, the speaker feature amount extraction section 133b extracts the first speaker feature amount from the acoustic feature amount using a DNN that has been subjected to learning.
[0077] In the present embodiment, the speaker feature amount extraction section 133b extracts the first speaker feature amount using an x-vector method, for example. Here, the x-vector method refers to a method of calculating a speaker feature amount that is a feature unique to a speaker called an x-Vector. More specifically, as shown in Figure 7 , the speaker feature amount extraction section 133b includes a frame concatenation processing section 1331 and a DNN 1332b.
[0078] [2.2.1.1-1 Frame concatenation processing section 1331]
[0079] The frame concatenation processing section 1331 performs the same processing as the frame concatenation processing section 121, that is, the frame concatenation processing section 1331 concatenates a plurality of frames of the MFCC that is the acoustic feature amount output from the preprocessing section 10 and outputs to the input layer of the DNN 1332b. In Figure 7 , the frame concatenation processing section 1331 outputs, to the input layer of the DNN 1332b, a 1200-dimensional vector generated by concatenating the MFCC parameters composed of 24-dimensional feature amounts per frame into 50 frames.
[0080] [2.2.1.1-2 DNN 1332b]
[0081] If a plurality of frames are input to the DNN 1332b from the frame concatenation processing section 1331, the DNN 1332b outputs the first speaker feature amount. In Figure 7In the example of FIG. 13B, the DNN 1332b is a neural network composed of an input layer, a plurality of intermediate layers, and an output layer, and is a neural network that has been learned using teacher voice data stored in the storage 1333b as teacher data. In the example of FIG. 13B, the DNN 1332b is a neural network that has been learned using teacher voice data composed of voices of a plurality of speakers, each of which includes the emotion "laugh". Figure 7 In the example of FIG. 13B, the storage 1333b stores teacher voice data composed of voices of a plurality of speakers, each of which includes the emotion "laugh".
[0082] In the example of FIG. 13B, the input layer is composed of, for example, 1200 nodes, and a 1200-dimensional vector generated by connecting MFCC parameters composed of 24-dimensional feature amounts per frame into 50 frames is input to the input layer. The output layer is composed of nodes that output speaker labels corresponding to the number of speakers included in the teacher data. In addition, the plurality of intermediate layers are composed of, for example, 2 to 3 layers, and the intermediate layer that calculates the 1st speaker feature amount is included therein. The intermediate layer that calculates the 1st speaker feature amount outputs the calculated 1st speaker feature amount as the output of the DNN 1332b. Figure 7 In the example of FIG. 13B, the input layer is composed of, for example, 1200 nodes, and a 1200-dimensional vector generated by connecting MFCC parameters composed of 24-dimensional feature amounts per frame into 50 frames is input to the input layer. The output layer is composed of nodes that output speaker labels corresponding to the number of speakers included in the teacher data. In addition, the plurality of intermediate layers are composed of, for example, 2 to 3 layers, and the intermediate layer that calculates the 1st speaker feature amount is included therein. The intermediate layer that calculates the 1st speaker feature amount outputs the calculated 1st speaker feature amount as the output of the DNN 1332b.
[0083] [2.2.1.2 Storage 134b]
[0084] The storage 134b is composed of, for example, a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores 2nd speaker feature amounts that are feature amounts unique to the registered speaker who is registered in advance and are feature amounts calculated from the registered speech of the registered speaker. In other words, the storage 134b stores 2nd speaker feature amounts that enable determination of each of voices that include the emotion of the registered speaker. More specifically, as illustrated in FIG. 13B, the storage 134b stores 2nd speaker feature amounts of the registered speech of the registered speaker that include the emotion "laugh" of the registered speaker. Figure 6
[0085] [2.2.1.3 Similarity Calculation Section 135b]
[0086] The similarity calculation section 135b calculates the similarity of the 1st speaker feature amount extracted by the speaker feature amount extraction section 133b and the 2nd speaker feature amount of the registered speaker stored in the storage 134b.
[0087] In the present embodiment, the similarity calculation section 135b calculates the similarity of each of the 1st speaker feature amount extracted by the speaker feature amount extraction section 133b and the 2nd speaker feature amount of one or more registered speakers stored in the storage 134b. The similarity calculation section 135b outputs a score indicating the calculated similarity.
[0088] For example, the similarity calculation unit 135b can also calculate the cosine using the inner product in the vector space model, and use the cosine distance (also called cosine similarity) of the angle between the vectors showing the features of the first speaker and the features of the second speaker as the similarity. In this case, the larger the value of the angle between the vectors, the lower the similarity. Alternatively, as a similarity calculation, the similarity calculation unit 135b can also use the inner product of the vector showing the features of the first speaker and the vector showing the features of the second speaker to calculate the cosine distance in the range of -1 to 1. In this case, the larger the value of the cosine distance, the higher the similarity.
[0089] Furthermore, since the speaker identifier 132a corresponding to "calm" and the speaker identifier 132b corresponding to "laugh" are the same, the description is omitted here.
[0090] [2.2.2 Recognizer Selection Unit 131]
[0091] The recognizer selection unit 131 selects a speaker recognizer 132k from a plurality of speaker recognizers 132 according to the emotion predicted by the emotion predictor 12. More specifically, the recognizer selection unit 131 selects a speaker recognizer 132k that has a second speaker feature stored in the storage unit, wherein the second speaker feature is a feature that can determine each of the voices containing an emotion of the registered speaker, and the emotion of the registered speaker corresponds to the emotion predicted by the emotion predictor 12. In addition, if there is no speaker recognizer 132 that corresponds to the emotion predicted by the emotion predictor 12, the recognizer selection unit 131 may also not use any speaker recognizer 132 (no selection).
[0092] Thus, the recognizer selection unit 131 can switch the speaker recognizer 132 based on the prediction result of the emotion inferrer 12.
[0093] [3. Speaker Identification Section 14]
[0094] like Figure 2 As shown, when the speaker recognition unit 14 is installed in the speaker recognition system 1, the speaker corresponding to the speech data is identified using the score output by the speaker recognition device 11.
[0095] In this embodiment, the speaker identification unit 14 identifies the speaker corresponding to the speech data based on a similarity score calculated by the similarity calculation unit 135b. For example, the speaker identification unit 14 uses such a score to output the registered speaker corresponding to the second speaker feature value that is closest to the first speaker feature value as the identification result.
[0096] [Operation of the speaker identification system 1]
[0097] Next, the operation of the speaker identification system 1 configured as above will be described.
[0098] Hereinafter, as the operation of the speaker identification system 1, the operation of the speaker identification device 11 having a characteristic operation will be described.
[0099] Figure 8 is a flowchart showing an outline of the operation of the speaker identification device 11 according to the present embodiment.
[0100] First, the speaker identification device 11 estimates the emotion contained in the speech sound shown by the speech data based on the acoustic feature quantity calculated from the speech data, using the DNN that has been learned (S11).
[0101] Next, the speaker identification device 11 outputs the score for identifying the speaker corresponding to the speech data based on the acoustic feature quantity calculated from the speech data, using the estimation result estimated in step S11 (S12).
[0102] (Effects, etc.)
[0103] As described above, with the speaker identification device 11 according to the present embodiment, the emotion estimator 12 that estimates the emotion of the evaluation speech is arranged at the front stage of the plurality of speaker identifiers 132 corresponding to each emotion, and the speaker identifiers 132 are switched according to the emotion shown by the estimation result of the emotion estimator 12.
[0104] Therefore, since the speaker identifier 132 corresponding to the emotion of the evaluation speech can be used, the speaker of the evaluation speech can be identified in a state where the emotion contained in the registration speech and the emotion contained in the evaluation speech are consistent.
[0105] Therefore, with the speaker identification device 11 according to the present embodiment, even if the emotion of the speaker is contained in the speech of the identification target, the identification accuracy of the speaker can be improved.
[0106] Further, with the speaker identification system 1 provided with the speaker identification device 11 according to the present embodiment, it is possible to identify the speaker of the sound in a free speech such as a conference recording system, a communication visualization system, and the like, that is, the sound in a conversation other than reading of an article and the like.
[0107] (Modified Example 1)
[0108] Further, as a method of identifying the speaker to whom the speech data showing the speech of the recognition target, i.e., the speech sound containing the speaker's emotion, corresponds, it is not limited to the above-described embodiment, i.e., it is not limited to the method in which the plurality of speaker identifiers 132 are constituted in the rear stage of the emotion estimator 12. Hereinafter, another method different from the method described in the above-described embodiment will be described as a modification example 1, and the difference from the above-described embodiment will be described as the center.
[0109] [4. Speaker identification apparatus 11A]
[0110] Figure 9 is a block diagram showing one example of a specific configuration of the speaker identification apparatus 11A related to the modification example 1 of the embodiment. Further, the same configuration as that of the speaker identification apparatus 11 shown in FIG. 1 is given the same symbol, and detailed description thereof will be omitted here. Figure 4
[0111] The speaker identification apparatus 11A performs processing for identifying the speaker to whom the speech data showing the speech sound of the recognition target corresponds. More specifically, the speaker identification apparatus 11A outputs, as the identification result, a score indicating the similarity of the first speaker characteristic amount evaluated for the speech to the third speaker characteristic amount which is the characteristic amount in which the second speaker characteristic amount of the registered speech of each registered speaker is modified.
[0112] As shown in FIG. 3, the speaker identification apparatus 11A related to the modification example has a different configuration of the speaker identification processing section 13A from that of the speaker identification apparatus 11 shown in FIG. 1. Figure 9 Figure 4
[0113] [4.1 Speaker identification processing section 13A]
[0114] The speaker identification processing section 13A outputs, using the estimation result of the emotion estimator 12, a score for identifying the speaker to whom the speech data corresponds, based on the acoustic characteristic amount calculated from the speech data.
[0115] As shown in FIG. 4, the speaker identification processing section 13A in the modification example has a speaker characteristic amount extraction section 133A, a storage section 134A, a similarity calculation section 135A, a storage section 136A, and a modification section 137A. Figure 9
[0116] [4.1.1 Speaker characteristic amount extraction section 133A]
[0117] The speaker feature amount extraction section 133A extracts the 1st speaker feature amount from the acoustic feature amount calculated from the speech data, the 1st speaker feature amount being a feature amount capable of identifying the speaker of the speech sound shown in the speech data.
[0118] In the present modification example, the speaker feature amount extraction section 133A extracts the 1st speaker feature amount using, for example, the x-vector method as well. For this reason, the speaker feature amount extraction section 133A can have a frame connection processing section and a DNN, like the speaker feature amount extraction section 133b. In the present modification example, learning is performed using, for example, teacher speech data composed of the sound of each of a plurality of speakers including "calm" as a recognition target. Note that "calm" is one example of an emotion, and can be another emotion such as "laugh". The description of other examples is omitted here because it has been described in the above-described embodiment.
[0119] [4.1.2 Storage Section 134A]
[0120] The storage section 134A is composed of, for example, a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores the 2nd speaker feature amount, which is a 2nd speaker feature amount registered in advance, and is a feature amount capable of identifying each of the sounds of the registered speaker including one emotion. As shown in FIG. 13, in the present modification example, the storage section 134A stores the 2nd speaker feature amount in the registered speech of the registered speaker including the emotion "calm". Note that "calm" is one example of an emotion, and can be another emotion such as "laugh". Figure 9
[0121] [4.1.3 Storage Section 136A]
[0122] The storage section 136A is composed of, for example, a rewritable nonvolatile memory such as a hard disk drive or a solid state drive, and stores learning data for modifying the emotion included in the registered speech. In the present modification example, the learning data stored in the storage section 136A is used to modify the 2nd speaker feature amount, which is the feature amount of the emotion "calm" stored in the storage section 134A, to the 3rd speaker feature amount, which is the speaker feature amount of the speech of the emotion corresponding to the emotion shown in the result of the speculation by the emotion speculator 12.
[0123] [4.1.4 Modification Section 137A]
[0124] The modification section 137A modifies the 2nd speaker feature amount stored in the storage section 134A to the 3rd speaker feature amount capable of identifying each of the sounds of the emotion corresponding to the emotion shown in the result of the speculation by the emotion speculator 12.
[0125] For example, the emotion indicated by the result of the speculation of the emotion speculator 12 is set to "laugh". In this case, the modification section 137A modifies the second speaker characteristic amount of the registered speech of the registered speaker, which contains the emotion "calm", stored in the storage section 134A, to a third speaker characteristic amount that can determine each of the sounds containing the emotion "laugh", for example, using the learning data stored in the storage section 136A. That is, the modification section 137A modifies the second speaker characteristic amount in the emotion "calm" stored in the storage section 134A to the third speaker characteristic amount in the emotion indicated by the result of the speculation of the emotion speculator 12, using the learning data stored in the storage section 136A.
[0126] [4.1.5 Similarity calculation section 135A]
[0127] The similarity calculation section 135A calculates the similarity of the first speaker characteristic amount extracted by the speaker characteristic amount extraction section 133A and the third speaker characteristic amount modified by the modification section 137A, and outputs the calculated similarity as a score.
[0128] In the present modification example, the similarity calculation section 135A calculates the similarity of the first speaker characteristic amount extracted by the speaker characteristic amount extraction section 133A and each of the third speaker characteristic amounts that are the characteristic amounts in which the second speaker characteristic amounts of one or more registered speakers stored in the storage section 134A are modified. The similarity calculation section 135A outputs a score indicating the calculated similarity.
[0129] [5. Speaker identification section 14]
[0130] The speaker identification section 14 identifies the speaker corresponding to the speech data using the score output by the speaker identification apparatus 11A.
[0131] In the present modification example, the speaker identification section 14 identifies the speaker corresponding to the speech data on the basis of the score indicated by the similarity calculated by the similarity calculation section 135A. For example, the speaker identification section 14 outputs, as an identification result, the registered speaker of the second speaker characteristic amount corresponding to the third speaker characteristic amount closest to the first speaker characteristic amount, using the score.
[0132] (Effects, etc.)
[0133] As described above, with the speaker identification apparatus 11A according to the present modification example, the speaker identification processing section 13A arranged in the latter stage identifies the speaker of the evaluation speech on the basis of the modification of the emotion of the registered speech to the emotion of the evaluation speech, in accordance with the result of the speculation of the emotion speculator 12 arranged in the former stage.
[0134] Thus, it is possible to recognize the speaker of the evaluation speech in a state where the emotion contained in the registration speech and the emotion contained in the evaluation speech are made consistent, that is, the difference in emotion, that is, the intonation, between the registration speech and the evaluation speech is modified to be consistent.
[0135] Therefore, by the speaker recognition device 11A according to the present modified example, it is possible to improve the recognition accuracy of the speaker even if the emotion of the speaker is contained in the speech of the recognition target.
[0136] (Modified Example 2)
[0137] The method explained in the above-described embodiment is not limited to the case explained in the embodiment and the modified example 1. Hereinafter, a different configuration of the speaker recognition device explained in the embodiment and the modified example 1 will be explained.
[0138] [6. Speaker recognition device 11B]
[0139] Figure 10 is a block diagram showing one example of the specific configuration of the speaker recognition device 11B according to the present modified example 2. Also, the same configuration as that of the speaker recognition device 11A shown in Figure 4 and Figure 9 will be given the same reference numerals, and detailed explanation thereof will be omitted here.
[0140] The speaker recognition device 11B, like the speaker recognition device 11, performs processing for recognizing the speaker to which the speech data showing the speech sound of the recognition target corresponds. More specifically, the speaker recognition device 11B calculates the similarity of the first speaker feature quantity of the evaluation speech and the second speaker feature quantity of the registration speech of each registration speaker. Also, the speaker recognition device 11B outputs, as the recognition result, the score obtained by giving reliability to the calculated similarity. In the present modified example, a case where weighting is performed to be used as the reliability will be explained.
[0141] As shown in Figure 10 , the speaker recognition device 11B according to the present modified example differs in the configuration of the speaker recognition processing section 13B from the speaker recognition device 11 shown in Figure 4 . Also, the speaker recognition device 11B according to the present modified example differs in the configuration of the speaker recognition processing section 13B from the speaker recognition device 11A shown in Figure 9 .
[0142] [6.1 Speaker recognition processing section 13B]
[0143] The speaker recognition processing section 13B outputs, as a score for identifying the speaker to whom the speech data corresponds, an acoustic feature quantity calculated from the speech data, using the result of the speculation by the emotion speculator 12.
[0144] Here, the acoustic feature quantity obtained by the speaker recognition processing section 13B is obtained by the preprocessing section 10 by dividing, in time series and by recognition units, all the speech data showing the speech sound of one speaker in a prescribed period, and thereby obtaining a plurality of speech data, the acoustic feature quantity being calculated from each of the obtained plurality of speech data,
[0145] In this modified example, as shown in Figure 10 The speaker recognition processing section 13B has a speaker feature quantity extraction section 133A, a storage section 134A, a similarity calculation section 135B, and a reliability assigning section 138B.
[0146] [6.1.1 Similarity Calculation Section 135B]
[0147] The similarity calculation section 135B calculates the similarity of the first speaker feature quantity extracted by the speaker feature quantity extraction section 133A and the second speaker feature quantity stored in the storage section 134A, and the second speaker feature quantity is a feature quantity that can determine each of the sounds containing one kind of emotion of the registered speaker as an object of recognition.
[0148] In this modified example, the similarity calculation section 135B calculates the similarity of the first speaker feature quantity extracted by the speaker feature quantity extraction section 133A and the second speaker feature quantity of the registered speaker containing the emotion of "calm" among the registered speakers of one or more stored in the storage section 134A.
[0149] [6.1.2 Reliability Assigning Section 138B]
[0150] The reliability assigning section 138B assigns, as a score, a weight corresponding to the emotion shown by the result of the speculation by the emotion speculator 12 to the similarity calculated by the similarity calculation section 135B, and outputs it. Here, in the case where one kind of emotion is consistent with the emotion shown by the result of the speculation, the reliability assigning section 138B assigns the greatest weight to the calculated similarity.
[0151] In this modified example, the reliability assigning section 138B assigns, to the similarity calculated by the similarity calculation section 135B for each of the plurality of speech data, a weight corresponding to the emotion shown by the result of the speculation for each of the plurality of speech data. The reliability assigning section 138B outputs, as a score for each of the plurality of speech data, the similarity weighted in each of the plurality of speech data to the speaker recognition section 14.
[0152] [7. The speaker identification unit 14]
[0153] As shown in FIG. 6, in a case where the speaker identification unit 14 is provided in the speaker identification system 1, the speaker corresponding to the speech data is identified using the score output by the speaker identification apparatus 11B. Figure 2
[0154] In the present modification example, the speaker identification unit 14 identifies the speaker corresponding to the speech data based on the score output by the similarity calculation unit 135B and indicating the weighted similarity. More specifically, the speaker identification unit 14 identifies the speaker corresponding to all the speech data using the total score, which is a score obtained by arithmetically averaging the scores output by the reliability assignment unit 138B for each of the plurality of speech data. Here, the speaker identification unit 14 identifies the speaker corresponding to the all speech data using the total score of the total scores that are equal to or higher than the threshold value. Also, the speaker identification unit 14 outputs the speaker identified as the result of identification corresponding to all the speech data. Thus, the speaker identification unit 14 can accurately identify the speaker corresponding to all the speech data corresponding to the total score using only the total score with high reliability.
[0155] [Example of the process of the speaker identification apparatus 11B]
[0156] Next, one example of the process of the speaker identification apparatus 11B configured as described above will be described. Figure 11
[0157] One example of the process of the speaker identification apparatus according to Modification Example 2 of the embodiment will be described. Figure 11 The first stage of FIG. 7 shows all the speech data obtained by the speaker identification apparatus 11B. Also, as described above, the all speech data refers to the sound signal of the sound converted from the speech sound of one speaker in a prescribed period, which is composed of the speech data divided by the identification unit. In the example of FIG. 7, the identification unit is an interval of 3 to 4 seconds, and the all speech data is the sound signal of the sound of 12 to 16 seconds, which is divided into the sound signal of 4 identification units. The all speech data divided by the identification unit corresponds to the speech data described above. Figure 11 Figure 11
[0158] Figure 11 The 2nd level of FIG. 2 shows the score before weighting and the inference result for each of the plurality of utterance data. The score before weighting indicates the similarity of each of the plurality of utterance data calculated by the speaker identification device 11B. The inference result is the emotion contained in the utterance sound indicated by each of the plurality of utterance data, which is inferred by the speaker identification device 11B for each of the plurality of utterance data constituting the entire utterance data. Figure 11 In the example of FIG. 2, the (score, emotion) is shown as (50, calm), (50, anger), (50, whisper), and (50, anger) for each of the plurality of utterance data (each of the utterance data) constituting the entire utterance data.
[0159] In addition, the 3rd level of FIG. 2 shows the score to which the weight is given on the basis of the inference result. The score indicates the similarity to which the weight is given on the basis of the inference result in each of the plurality of utterance data, that is, indicates the similarity in each of the plurality of utterance data. In the example of FIG. 2, the weight is given as 75, 25, 5, and 25 for each of the plurality of utterance data (each of the utterance data) constituting the entire utterance data, when the emotion indicated by the inference result is "calm". In addition, the weight is given as the maximum when the emotion indicated by the inference result is "calm". This is because the speaker identification device 11B calculates the similarity in each of the plurality of utterance data using the 2nd speaker characteristic amount of the registered utterance in which the emotion "calm" of the registered speaker is contained. That is, the more consistent the emotion contained in the registered utterance used for calculating the similarity in each of the plurality of utterance data by the speaker identification device 11B is, the higher the reliability of the calculated similarity is set, and the larger the weight is given. Figure 11 Figure 11 The 4th level of FIG. 2 shows the total score. The total score is the score for the entire utterance data, which is the average of the scores for each of the plurality of utterance data as described above. In the example of FIG. 2, the calculated value is 32.5.
[0160] Figure 11 Figure 11 (Effects and the like)
[0161] (Effects and the like)
[0162] As described above, in the speaker identification device 11B according to the present modified example, the speaker identification processing section 13B outputs the score obtained by giving the weight based on the inference result of the emotion of the evaluation utterance to the similarity calculated for the evaluation utterance and the registered utterance. In addition, the more consistent the emotion contained in the evaluation utterance indicated by the inference result is with the emotion contained in the registered utterance, the higher the reliability of the calculated similarity is set, and the larger the weight is given.
[0163] Accordingly, it is possible to identify the speaker of the evaluation speech in a state in which the emotion included in the evaluation speech is close (similar) to the emotion included in the registered speech, using the score with high reliability.
[0164] Accordingly, by the speaker identification device 11B according to the present modification, it is possible to improve the accuracy of speaker identification even if the emotion of the speaker is included in the speech of the identification target.
[0165] In addition, it is also possible to confirm the reliability of the result of speaker identification by confirming the reliability of the score.
[0166] (Modification 3)
[0167] In the case described in Modification 2, the speaker identification device 11B outputs the score obtained by giving reliability to the calculated similarity, and the reliability given here is the weight calculated based on the estimation result of the emotion included in the evaluation speech. In Modification 3, the speaker identification device 11C gives reliability to the calculated similarity and outputs it, and the reliability given here is the reliability (specifically, additional information indicating the reliability) based on the estimation result of the emotion included in the evaluation speech. Hereinafter, the speaker identification device 11C according to Modification 3 will be described focusing on the differences from the speaker identification device 11B described in Modification 2.
[0168] [8. Speaker identification device 11C]
[0169] Figure 12 is a block diagram showing one example of the specific configuration of the speaker identification device 11C according to Modification 3 of the present embodiment. In addition, the same symbols are given to the same configurations as those of the speaker identification device 11B described in Modification 2, and detailed description thereof will be omitted here. Figure 4 , Figure 9 and Figure 10 and the like, and the detailed description thereof will be omitted here.
[0170] The speaker identification device 11C, like the speaker identification device 11B, performs processing for identifying the speaker of the speech data showing the speech sound of the identification target. More specifically, the speaker identification device 11C calculates the similarity of the first speaker feature amount of the evaluation speech and the second speaker feature amount of the registered speech of each registered speaker. Then, the speaker identification device 11B outputs, as the identification result, the score obtained by giving reliability (or additional information indicating the reliability) to the calculated similarity.
[0171] As shown in Figure 12 , the speaker identification device 11C according to the present modification gives reliability to the calculated similarity with respect to Figure 10The speaker identification device 11B is different from the speaker identification device 11A in that the configuration of the speaker identification processing section 13C is different. More specifically, the speaker identification device 11C according to the present modified example is different from the speaker identification device 11B in that Figure 10 The speaker identification device 11B is different from the speaker identification device 11A in that the configuration of the speaker identification processing section 13C is different. More specifically, the speaker identification device 11C according to the present modified example is different from the speaker identification device 11B in that
[0172] [8.1 Reliability-imparting section 138C]
[0173] The reliability-imparting section 138C imparts reliability corresponding to the emotion indicated by the estimation result of the emotion estimator 12 to the similarity calculated by the similarity calculation section 135B, and outputs the result as a score. Here, the reliability-imparting section 138C imparts the highest reliability to the calculated similarity in a case where one emotion is consistent with the emotion indicated by the estimation result.
[0174] [9. Speaker identification section 14]
[0175] The speaker identification section 14 identifies the speaker corresponding to the speech data using the score output by the speaker identification device 11C.
[0176] In the present modified example, the speaker identification section 14 identifies the speaker corresponding to the speech data based on the score indicating the similarity to which reliability is imparted, which is output by the similarity calculation section 135B. For example, the speaker identification section 14 identifies the speaker corresponding to the speech data using the score to which reliability of a threshold value or more is imparted. Also, the speaker identification section 14 outputs the speaker identified as the result of identification. According to this, the speaker identification section 14 can accurately identify the speaker corresponding to the speech data corresponding to the score using only the score with high reliability.
[0177] (Effects, etc.)
[0178] As described above, in the speaker identification device 11C according to the present modified example, the speaker identification processing section 13C outputs a score obtained by imparting additional information based on the estimation result of the emotion evaluating the speech to the similarity calculated with respect to the evaluation speech and the registered speech. For example, the speaker identification processing section 13C imparts the additional information so that the reliability with respect to the calculated similarity becomes higher as the emotion included in the evaluation speech indicated by the estimation result is more consistent with the emotion included in the registered speech.
[0179] According to this, it is possible to identify the speaker of the evaluation speech in a state where the emotion included in the registered speech is close (similar) to the emotion included in the evaluation speech using the score with high reliability.
[0180] Thus, with the speaker identification device 11C according to the present modified example, even if the speaker's emotions are included in the speech of the identification target, it is possible to improve the accuracy of speaker identification.
[0181] In addition, the reliability of the result of speaker identification can be confirmed by confirming the reliability of the score.
[0182] (Possibilities of Other Embodiments)
[0183] The above describes the speaker identification device according to the embodiments and the modified examples, but the present disclosure is not limited to these embodiments.
[0184] For example, each processing section included in the speaker identification device according to the above embodiments and the modified examples can be implemented as a typical integrated circuit, i.e., LSI. These can be individually made into one chip, or some or all of them can be made into one chip.
[0185] In addition, the method of integration is not limited to LSI, and can be implemented by a dedicated circuit or a general-purpose processor. After LSI manufacture, a programmable FPGA (Field Programmable Gate Array) or a reconfigurable processor that can reconfigure the connection or settings of circuit cells inside the LSI can be used.
[0186] In addition, the present disclosure can be implemented as a speaker identification method executed by the speaker identification device.
[0187] In addition, each constituent element in each of the above embodiments can be constituted by a dedicated hardware or implemented by executing a software program suitable for each constituent element. Each constituent element can be implemented by a program execution section such as a CPU or a processor reading and executing a software program recorded in a storage medium such as a hardware or a semiconductor memory.
[0188] In addition, the division of the functional blocks in the block diagram is one example, and a plurality of functional blocks can be implemented as one functional block, one functional block can be divided into a plurality of functional blocks, and part of the functions can be transferred to other functional blocks. Also, the functions of a plurality of functional blocks having similar functions can be processed in parallel or time-division by a single hardware or software.
[0189] In addition, the order of the steps performed in the flowchart is for the purpose of specifically describing an example of the present disclosure, and can be in an order other than the above. Also, part of the above steps can be executed simultaneously (in parallel) with other steps.
[0190] Although the speaker identification device according to one or more aspects of the present disclosure has been described above based on embodiments and modifications, the present disclosure is not limited to these embodiments and modifications. Various modifications that can be conceived by those skilled in the art, modes obtained by applying the modifications to the respective embodiments, and modes obtained by combining the components of different embodiments and modifications are included in one or more aspects of the present disclosure.
[0191] The present disclosure can be used for a speaker identification device, a speaker identification method, and a recording medium, and can identify a speaker of a free speech including emotion, such as a conference recording system, a communication visualization system, using the speaker identification device, the speaker identification method, and the recording medium.
[0192] Symbol explanation
[0193] 1 Speaker identification system
[0194] 10 Preprocessing section
[0195] 11, 11A, 11B, 11C Speaker identification device
[0196] 12 Emotion estimator
[0197] 13, 13A, 13B, 13C Speaker identification processing section
[0198] 14 Speaker identification section
[0199] 101 Sound acquisition section
[0200] 102 Acoustic feature amount calculation section
[0201] 121, 1331 Frame connection processing section
[0202] 122, 1332b DNN
[0203] 123, 134A, 134B, 136A, 1333b Storage section
[0204] 131 Identifier selection section
[0205] 132, 132A, 132B Speaker identifier
[0206] 133A, 133B Speaker feature amount extraction section
[0207] 135A, 135B, 135B Similarity calculation section
[0208] 137A Modification section
[0209] 138B, 138C reliability assignment unit
Claims
1. A speaker recognition device for identifying the speaker corresponding to speech data showing the speech voice of the identified object. The speaker recognition device includes: An emotion inferrer, utilizing a learned deep neural network, infers the emotion contained in the spoken voice as shown in the speech data based on acoustic feature quantities calculated from the speech data; and The speaker recognition processing unit extracts a first speaker feature quantity for evaluating the speech based on the acoustic feature quantity calculated from the speech data, calculates the similarity between the second speaker feature quantity selected from the second speaker feature quantities pre-registered by each registered speaker and the extracted first speaker feature quantity, which corresponds to the emotion indicated by the emotion inference result of the emotion inferrer, and outputs a score for identifying the speaker corresponding to the speech data based on the similarity.
2. A speaker recognition device for identifying the speaker corresponding to speech data showing the speech voice of the identified object. The speaker recognition device includes: An emotion inferrer, utilizing a learned deep neural network, infers the emotion contained in the spoken voice as shown in the speech data based on acoustic feature quantities calculated from the speech data; and The speaker recognition processing unit, using the inference results of the emotion inferrer and based on the acoustic feature quantities calculated from the speech data, outputs a score for identifying the speaker corresponding to the speech data. The speaker recognition processing unit includes multiple speaker recognizers and a recognizer selection unit. Each of the plurality of speaker recognizers includes a speaker feature extraction unit and a similarity calculation unit. The speaker feature extraction unit, when acoustic features are input, extracts a first speaker feature from the input acoustic features. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity calculation unit calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit and a second speaker feature stored in a storage unit. This second speaker feature is a feature capable of identifying each voice containing an emotion of the registered speaker being identified. The recognizer selection unit selects one speaker recognizer from the plurality of speaker recognizers. The selected speaker recognizer is one that stores in the storage unit a second speaker characteristic quantity that can determine each of the following sounds: each sound is a sound that corresponds to the emotion indicated by the inferred result and contains an emotion of the registered speaker. The speaker recognizer selected by the recognizer selection unit calculates the similarity by using acoustic feature quantities calculated from the speech data as input, and outputs it as the score.
3. A speaker recognition device for identifying the speaker corresponding to speech data showing the speech voice of the identified object. The speaker recognition device includes: An emotion inferrer uses a learned deep neural network to infer the emotions contained in the spoken voice as shown in the speech data, based on acoustic feature quantities calculated from the speech data. as well as The speaker recognition processing unit, using the inference results of the emotion inferrer and based on the acoustic feature quantities calculated from the speech data, outputs a score for identifying the speaker corresponding to the speech data. The speaker recognition and processing unit has the following capabilities: The speaker feature extraction unit extracts a first speaker feature from the acoustic features, which can determine the speaker of the speech voice shown in the speech data; The modification unit modifies the second speaker feature stored in the storage unit into a third speaker feature. The second speaker feature can determine each of the voices containing an emotion of the registered speaker as the object of identification, and the third speaker feature can determine each of the voices containing the emotion corresponding to the emotion shown in the inferred result. as well as The similarity calculation unit calculates the similarity between the extracted first speaker features and the third speaker features modified by the modification unit, and outputs the calculated similarity as the score.
4. A speaker recognition device for identifying the speaker corresponding to speech data showing the speech voice of the identified object. The speaker recognition device includes: An emotion inferrer uses a learned deep neural network to infer the emotions contained in the spoken voice as shown in the speech data, based on acoustic feature quantities calculated from the speech data. as well as The speaker recognition processing unit, using the inference results of the emotion inferrer and based on the acoustic feature quantities calculated from the speech data, outputs a score for identifying the speaker corresponding to the speech data. The speaker recognition and processing unit has the following capabilities: The speaker feature extraction unit extracts a first speaker feature from the acoustic features, which can determine the speaker of the speech voice shown in the speech data; The similarity calculation unit calculates the similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. as well as The reliability assignment unit assigns a weight corresponding to the sentiment indicated by the inferred result to the calculated similarity, and outputs it as the score. The reliability assignment unit assigns the maximum weight to the calculated similarity when the emotion matches the emotion shown in the inferred result.
5. The speaker recognition device as described in claim 4, The acoustic features are calculated as follows: the preprocessing unit segments all speech data of a speaker's voice within a specified period, in a time sequence and according to recognition units, to obtain multiple speech data sets. The acoustic features are calculated from each of the obtained multiple speech data sets. The reliability assignment unit assigns weights to the similarity and outputs it as the score. The similarity is calculated by the similarity calculation unit for each of the plurality of speech data. The weights are weights corresponding to the emotions inferred by the emotion inferr for each of the plurality of speech data.
6. The speaker recognition device as described in claim 5, The speaker recognition device further includes a speaker recognition unit that identifies the speaker corresponding to all the speech data using an overall score, wherein the overall score is obtained by arithmetically averaging the scores output by the reliability assignment unit for each of the plurality of speech data. The speaker identification unit uses the overall score above a threshold in the overall score to identify the speaker corresponding to all the speech data.
7. A speaker recognition device for identifying the speaker corresponding to speech data showing the speech voice of the identified object. The speaker recognition device includes: An emotion inferrer uses a learned deep neural network to infer the emotions contained in the spoken voice as shown in the speech data, based on acoustic feature quantities calculated from the speech data. as well as The speaker recognition processing unit, using the inference results of the emotion inferrer and based on the acoustic feature quantities calculated from the speech data, outputs a score for identifying the speaker corresponding to the speech data. The speaker recognition and processing unit has the following capabilities: The speaker feature extraction unit extracts a first speaker feature from the acoustic features, which can determine the speaker of the speech voice shown in the speech data; The similarity calculation unit calculates the similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. as well as The reliability assignment unit assigns a reliability corresponding to the sentiment indicated by the inference result to the calculated similarity and outputs it as the score.
8. The speaker recognition device as described in claim 7, The speaker recognition device further includes a speaker recognition unit, which uses the score, which is above the reliability threshold, to identify the speaker corresponding to the speech data.
9. The speaker recognition device as claimed in any one of claims 2 to 8, The speaker feature extraction unit uses a learned deep neural network to extract the first speaker features from the acoustic features.
10. A speaker recognition method, comprising identifying the speaker corresponding to speech data showing the speech voice of an identified object, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step involves extracting a first speaker feature quantity for evaluating the speech based on the acoustic feature quantity calculated from the speech data, calculating the similarity between the second speaker feature quantity selected from the second speaker feature quantities pre-registered by each registered speaker and the extracted first speaker feature quantity, which corresponds to the emotion indicated by the prediction result in the emotion inference step, and outputting a score for identifying the speaker corresponding to the speech data based on the similarity.
11. A speaker recognition method, comprising identifying the speaker corresponding to speech data showing the speech voice of an identified object, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A speaker recognizer is selected from a plurality of speaker recognizers, wherein the selected speaker recognizer is one that stores in its storage a second speaker characteristic quantity capable of determining each of the following sounds: each of the sounds is a sound containing an emotion of the registered speaker being identified, corresponding to the emotion indicated by the inferred result. Each of the plurality of speaker recognizers includes a speaker feature extraction unit and a similarity calculation unit. The speaker feature extraction unit, when acoustic features are input, extracts a first speaker feature from the input acoustic features. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity calculation unit calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit and a second speaker feature stored in a storage unit. This second speaker feature is a feature capable of identifying each voice containing an emotion of the registered speaker being identified. The selected speaker recognizer calculates the similarity by taking acoustic feature quantities calculated from the speech data as input, and outputs the similarity as the score.
12. A speaker recognition method, comprising identifying the speaker corresponding to speech data showing the speech voice of an identified object, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The second speaker feature stored in the storage unit is modified into a third speaker feature. The second speaker feature can determine each voice containing an emotion of the registered speaker as the identification object, and the third speaker feature can determine each voice containing the emotion corresponding to the emotion shown in the inferred result. Calculate the similarity between the extracted first speaker features and the modified third speaker features, and output the calculated similarity as the score.
13. A speaker recognition method, comprising identifying the speaker corresponding to speech data showing the speech voice of an identified object, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit is calculated, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. The calculated similarity is assigned a weight corresponding to the sentiment indicated by the inferred result, and this weight is output as the score. When the emotion in question matches the emotion indicated by the inferred result, the calculated similarity is assigned the highest weight.
14. A speaker recognition method, comprising identifying the speaker corresponding to speech data showing the speech voice of an identified object, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit is calculated, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. The calculated similarity is assigned a reliability corresponding to the sentiment indicated by the inferred result, and output as the score.
15. A recording medium, a computer-readable non-transitory recording medium, wherein a program is recorded in the recording medium, the program causing a computer to execute a speaker recognition method, the speaker recognition method identifying a speaker corresponding to speech data showing the speech voice of an object to be identified, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step involves extracting a first speaker feature quantity for evaluating the speech based on the acoustic feature quantity calculated from the speech data, calculating the similarity between the selected second speaker feature quantity corresponding to the emotion indicated by the inference result in the emotion inference step from the second speaker feature quantity pre-registered by each registered speaker, and the extracted first speaker feature quantity, and outputting a score for identifying the speaker corresponding to the speech data based on the similarity.
16. A recording medium, a computer-readable non-transitory recording medium, wherein a program is recorded in the recording medium, the program causing a computer to execute a speaker recognition method, the speaker recognition method identifying a speaker corresponding to speech data showing the speech voice of an object to be identified, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A speaker recognizer is selected from a plurality of speaker recognizers, wherein the selected speaker recognizer is one that stores in its storage a second speaker characteristic quantity capable of determining each of the following sounds: each of the sounds is a sound containing an emotion of the registered speaker being identified, corresponding to the emotion indicated by the inferred result. Each of the plurality of speaker recognizers includes a speaker feature extraction unit and a similarity calculation unit. The speaker feature extraction unit, when acoustic features are input, extracts a first speaker feature from the input acoustic features. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity calculation unit calculates the similarity between the first speaker feature extracted by the speaker feature extraction unit and a second speaker feature stored in a storage unit. This second speaker feature is a feature capable of identifying each voice containing an emotion of the registered speaker being identified. The selected speaker recognizer calculates the similarity by taking acoustic feature quantities calculated from the speech data as input, and outputs the similarity as the score.
17. A recording medium, a computer-readable non-transitory recording medium, wherein a program is recorded in the recording medium, the program causing a computer to execute a speaker recognition method, the speaker recognition method identifying a speaker corresponding to speech data showing the speech voice of an object to be identified, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The second speaker feature stored in the storage unit is modified into a third speaker feature. The second speaker feature can determine each voice containing an emotion of the registered speaker as the identification object, and the third speaker feature can determine each voice containing the emotion corresponding to the emotion shown in the inferred result. Calculate the similarity between the extracted first speaker features and the modified third speaker features, and output the calculated similarity as the score.
18. A recording medium, a computer-readable non-transitory recording medium, wherein a program is recorded in the recording medium, the program causing a computer to execute a speaker recognition method, the speaker recognition method identifying a speaker corresponding to speech data showing the speech voice of an object to be identified, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit is calculated, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. The calculated similarity is assigned a weight corresponding to the sentiment indicated by the inferred result, and this weight is output as the score. When the emotion in question matches the emotion indicated by the inferred result, the calculated similarity is assigned the highest weight.
19. A recording medium, a computer-readable non-transitory recording medium, wherein a program is recorded in the recording medium, the program causing a computer to execute a speaker recognition method, the speaker recognition method identifying a speaker corresponding to speech data showing the speech voice of an object to be identified, the speaker recognition method comprising the following steps: The emotion inference step utilizes a learned deep neural network to infer the emotion contained in the spoken voice based on acoustic feature quantities calculated from the speech data; and The speaker identification processing step utilizes the inference results from the emotion inference step and outputs a score for identifying the speaker corresponding to the speech data based on the acoustic feature quantities calculated from the speech data. In the speaker identification processing step, A first speaker feature is extracted from the acoustic feature quantities. This first speaker feature is capable of identifying the speaker of the speech voice shown in the speech data. The similarity between the extracted first speaker feature and the second speaker feature stored in the storage unit is calculated, and the second speaker feature is a feature that can be determined for each of the voices containing an emotion of the registered speaker as the object of identification. The calculated similarity is assigned a reliability corresponding to the sentiment indicated by the inferred result, and output as the score.
Citation Information
Patent Citations
Registered utterance division device, speaker likelihood evaluation device, speaker identification device, registered utterance division method, speaker likelihood evaluation method, and program
JP2017187642A
Method of recognizing talker and device therefor
JP1998247092A