Speaker recognition method, speaker recognition device, and speaker recognition program

The proposed method improves speaker identification accuracy by selecting gender-specific models based on the speaker's gender, effectively addressing the variability in voice data feature distributions between males and females.

JP7696331B2Active Publication Date: 2025-06-20PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022509394
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2021-02-15
Publication Date
2025-06-20
Estimated Expiration
2041-02-15

AI Technical Summary

Technical Problem

Conventional speaker identification techniques struggle to improve the accuracy of identifying whether a speaker is pre-registered, especially due to differences in voice data feature distributions between males and females.

Method used

A speaker identification method that acquires and processes voice data by selecting either a male-specific or female-specific speaker identification model based on the gender of the speaker, improving accuracy by matching feature amounts with gender-specific models.

Benefits of technology

Enhances the accuracy of speaker identification by utilizing gender-specific models, effectively addressing the variability in voice data feature distributions between males and females.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696331000001
    Figure 0007696331000001
  • Figure 0007696331000002
    Figure 0007696331000002
  • Figure 0007696331000003
    Figure 0007696331000003
Patent Text Reader

Abstract

This speaker identification device: acquires audio data to be identified; acquires registered audio data; selects a first speaker identification model that has undergone machine learning by using man's audio data in order to identify a male speaker if the sex of either the speaker of the audio data to be identified or the speaker of the registered audio data is male; selects a second speaker identification model that has undergone machine learning by using woman's audio data in order to identify a female speaker if the sex of either the speaker of the audio data to be identified or the speaker of the registered audio data is female; and inputs a feature amount of the audio data to be identified and a feature amount of the registered audio data to the first speaker identification model or the second speaker identification model which was selected, thereby identifying the speaker of the audio data subject to identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technique for identifying a speaker.

Background Art

[0002] Conventionally, a technique is known in which voice data of a speaker to be identified is acquired, and based on the acquired voice data, it is identified whether the speaker to be identified is a pre-registered speaker. In conventional speaker identification, the similarity between the feature amount of the voice data of the speaker to be identified and the feature amount of the voice data of the registered speaker is calculated. If the calculated similarity is equal to or greater than a threshold value, it is determined that the speaker to be identified and the registered speaker are the same.

[0003] For example, Non-Patent Document 1 discloses a speaker-specific feature amount called i-vector as a high-precision feature amount for speaker identification.

[0004] Also, for example, Non-Patent Document 2 discloses an x-vector as a feature amount that replaces the i-vector. The x-vector is a feature amount extracted by inputting voice data into a deep neural network generated by deep learning.

[0005] However, in the above conventional technology, further improvement has been required to improve the accuracy of identifying whether the speaker to be identified is a pre-registered speaker.

Prior Art Documents

Non-Patent Documents

[0006]

Non-Patent Document 1

[0007] The present disclosure has been made to solve the above problems, and an object thereof is to provide a technique capable of improving the accuracy of identifying whether or not a speaker to be identified is a pre-registered speaker.

[0008] A speaker identification method according to an aspect of the present disclosure includes a computer acquiring identification target voice data, acquiring pre-registered registered voice data, extracting a feature amount of the identification target voice data, extracting a feature amount of the registered voice data, and when either the speaker of the identification target voice data or the speaker of the registered voice data is male, selecting a first speaker identification model learned by machine learning using male voice data to identify a male speaker, and when either the speaker of the identification target voice data or the speaker of the registered voice data is female, selecting a second speaker identification model learned by machine learning using female voice data to identify a female speaker, and inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, thereby identifying the speaker of the identification target voice data.

[0009] According to the present disclosure, it is possible to improve the accuracy of identifying whether a speaker to be identified is a pre-registered speaker or not.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Embodiments for Carrying Out the Invention

[0011] (Findings underlying the present disclosure) The distributions of the feature amounts of the uttered voice data are different between men and women. However, conventional speaker identification devices identify the speaker of the voice data to be identified using a speaker identification model generated from voice data collected regardless of the gender of the speaker. Thus, in the prior art, speaker identification specialized for each of male speakers and female speakers has not been considered. Therefore, there is a possibility that the accuracy of speaker identification can be improved by performing speaker identification considering gender.

[0012] In order to solve the above problems, a speaker identification method according to one aspect of the present disclosure includes a computer acquiring identification target voice data, acquiring registered voice data registered in advance, extracting a feature amount of the identification target voice data, extracting a feature amount of the registered voice data, and when either the speaker of the identification target voice data or the speaker of the registered voice data is male, selecting a first speaker identification model that has been machine-learned using male voice data to identify a male speaker, and when either the speaker of the identification target voice data or the speaker of the registered voice data is female, selecting a second speaker identification model that has been machine-learned using female voice data to identify a female speaker, and inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model to identify the speaker of the identification target voice data.

[0013] According to this configuration, when either the speaker of the identification target voice data or the speaker of the registered voice data is male, the feature amount of the identification target voice data and the feature amount of the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the identification target voice data is identified. Also, when either the speaker of the identification target voice data or the speaker of the registered voice data is female, the feature amount of the identification target voice data and the feature amount of the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the identification target voice data is identified.

[0014] Therefore, even when the distribution of the feature amounts of the voice data differs depending on gender, the speaker of the identification target voice data is identified by the first speaker identification model and the second speaker identification model specialized for each gender, so that the accuracy of identifying whether the speaker to be identified is a speaker registered in advance can be improved.

[0015] Also, in the above speaker identification method, when selecting the first speaker identification model or the second speaker identification model, if the gender of the speaker in the registered voice data is male, the first speaker identification model may be selected, and if the gender of the speaker in the registered voice data is female, the second speaker identification model may be selected.

[0016] According to this configuration, either the first speaker identification model or the second speaker identification model is selected according to the gender of the speaker in the registered voice data. Therefore, by identifying the gender of the speaker in the registered voice data only once at the time of registration, it is not necessary to identify the gender of the voice data to be identified for each speaker identification, and the processing load of speaker identification can be reduced.

[0017] Also, in the above speaker identification method, further, acquisition target voice data is acquired, further, feature amounts of the acquisition target voice data are extracted, further, the gender of the speaker of the acquisition target voice data is identified using the feature amounts of the acquisition target voice data, and further, the acquisition target voice data associated with the identified gender may be registered as the registered voice data.

[0018] According to this configuration, at the time of registration, registered voice data associated with gender can be registered in advance, and at the time of speaker identification, either the first speaker identification model or the second speaker identification model can be easily selected using the gender associated with the registered voice data registered in advance.

[0019] Also, in the above speaker identification method, in the identification of the gender, a gender identification model learned by machine learning using voice data of male and female may be acquired to identify the gender of the speaker, and the feature amounts of the acquisition target voice data are input to the gender identification model, whereby the gender of the speaker of the acquisition target voice data may be identified.

[0020] According to this configuration, just by inputting the voice data to be registered into the gender identification model that has been machine-learned using male and female voice data in order to identify the gender of the speaker, it is possible to easily identify the gender of the speaker of the voice data to be registered.

[0021] Also, in the above speaker identification method, in the identification of the gender, by inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored male voice data into the gender identification model, the similarity between the voice data to be registered and each of the plurality of male voice data is obtained from the gender identification model, and the average of the obtained plurality of similarities is calculated as the average male similarity. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored female voice data into the gender identification model, the similarity between the voice data to be registered and each of the plurality of female voice data is obtained from the gender identification model, and the average of the obtained plurality of similarities is calculated as the average female similarity. When the average male similarity is higher than the average female similarity, the gender of the speaker of the voice data to be registered may be identified as male. When the average male similarity is lower than the average female similarity, the gender of the speaker of the voice data to be registered may be identified as female.

[0022] According to this configuration, when the speaker of the voice data to be registered is male, the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data is higher than the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of female voice data. Also, when the speaker of the voice data to be registered is female, the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of female voice data is higher than the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data. Therefore, by comparing the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data with the average similarity between the feature amount of the voice data to be registered and the feature amounts of a plurality of female voice data, it is possible to easily identify the gender of the speaker of the voice data to be registered.

[0023] Also, in the above speaker identification method, in the identification of the gender, by inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored male voice data into the gender identification model, the similarity between the voice data to be registered and each of the plurality of male voice data is obtained from the gender identification model, and the maximum value among the obtained plurality of similarities is calculated as the maximum male similarity. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored female voice data into the gender identification model, the similarity between the voice data to be registered and each of the plurality of female voice data is obtained from the gender identification model, and the maximum value among the obtained plurality of similarities is calculated as the maximum female similarity. When the maximum male similarity is higher than the maximum female similarity, the gender of the speaker of the voice data to be registered may be identified as male. When the maximum male similarity is lower than the maximum female similarity, the gender of the speaker of the voice data to be registered may be identified as female.

[0024] According to this configuration, when the speaker of the voice data to be registered is male, the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of male voice data is higher than the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of female voice data. Also, when the speaker of the voice data to be registered is female, the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of female voice data is higher than the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of male voice data. Therefore, by comparing the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of male voice data with the maximum similarity among the plurality of similarities between the feature amount of the voice data to be registered and the feature amounts of the plurality of female voice data, the gender of the speaker of the voice data to be registered can be easily identified.

[0025] Also, in the above speaker identification method, in the identification of the gender, the average feature amount of a plurality of pre-stored male voice data is calculated, and the feature amount of the registration target voice data and the average feature amount of the plurality of male voice data are input into the gender identification model, so as to obtain a first similarity between the registration target voice data and the male voice data group from the gender identification model. The average feature amount of a plurality of pre-stored female voice data is calculated, and the feature amount of the registration target voice data and the average feature amount of the plurality of female voice data are input into the gender identification model, so as to obtain a second similarity between the registration target voice data and the female voice data group from the gender identification model. When the first similarity is higher than the second similarity, the gender of the speaker of the registration target voice data is identified as male. When the first similarity is lower than the second similarity, the gender of the speaker of the registration target voice data may be identified as female.

[0026] According to this configuration, when the speaker of the registration target voice data is male, the first similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of male voice data is higher than the second similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of female voice data. Also, when the speaker of the registration target voice data is female, the second similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of female voice data is higher than the first similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of male voice data. Therefore, by comparing the first similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of male voice data with the second similarity between the feature amount of the registration target voice data and the average feature amount of the plurality of female voice data, the gender of the speaker of the registration target voice data can be easily identified.

[0027] Further, in the above speaker identification method, the registered voice data includes a plurality of registered voice data. In identifying the speaker, the feature amount of the voice data to be identified and the feature amounts of the plurality of registered voice data are input into either the selected first speaker identification model or the second speaker identification model, so as to obtain the similarity between the voice data to be identified and each of the plurality of registered voice data from either the first speaker identification model or the second speaker identification model. The speaker of the registered voice data with the highest obtained similarity may be identified as the speaker of the voice data to be identified.

[0028] According to this configuration, the similarity between the voice data to be identified and each of the plurality of registered voice data is obtained from either the first speaker identification model or the second speaker identification model, and the speaker of the registered voice data with the highest similarity is identified as the speaker of the voice data to be identified. Therefore, among the plurality of registered voice data, the speaker of the most similar registered voice data can be identified as the speaker of the voice data to be identified.

[0029] Further, in the above speaker identification method, the registered voice data includes a plurality of registered voice data, and the plurality of registered voice data is associated with identification information for identifying the speaker of each of the plurality of registered voice data. Further, identification information for identifying the speaker of the voice data to be identified is obtained. In obtaining the registered voice data, the registered voice data associated with the identification information that matches the obtained identification information is obtained from among the plurality of registered voice data. In identifying the speaker, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into either the selected first speaker identification model or the second speaker identification model, so as to obtain the similarity between the voice data to be identified and the registered voice data from either the first speaker identification model or the second speaker identification model. When the obtained similarity is higher than a threshold value, the speaker of the registered voice data may be identified as the speaker of the voice data to be identified.

[0030] According to this configuration, among a plurality of registered voice data, one piece of registered voice data associated with identification information that matches the identification information for identifying the speaker of the voice data to be identified is acquired. Therefore, it is not necessary to calculate the similarity between the feature amounts of all the registered voice data among the plurality of registered voice data and the feature amount of the voice data to be identified. Instead, it is only necessary to calculate the similarity between the feature amount of one piece of registered voice data among the plurality of registered voice data and the feature amount of the voice data to be identified. Thus, the processing load of speaker identification can be reduced.

[0031] Also, in the above-described speaker identification method, during machine learning, all combinations of the feature amounts of two pieces of voice data among a plurality of male voice data are input into the first speaker identification model, whereby the similarity of each of the plurality of combinations of the two pieces of voice data is obtained from the first speaker identification model, and a first threshold value capable of distinguishing the similarity of the two pieces of voice data of the same speaker from the similarity of the two pieces of voice data of different speakers is calculated. During machine learning, all combinations of the feature amounts of two pieces of voice data among a plurality of female voice data are input into the second speaker identification model, whereby the similarity of each of the plurality of combinations of the two pieces of voice data is obtained from the second speaker identification model, and a second threshold value capable of distinguishing the similarity of the two pieces of voice data of the same speaker from the similarity of the two pieces of voice data of different speakers is calculated. In the identification of the speaker, when the similarity is obtained from the first speaker identification model, the first threshold value may be subtracted from the obtained similarity, and when the similarity is obtained from the second speaker identification model, the second threshold value may be subtracted from the obtained similarity.

[0032] When two different first speaker identification models and a second speaker identification model are used, the ranges of the output values of the first speaker identification model and the second speaker identification model may be different. Therefore, at the time of registration, for each of the first speaker identification model and the second speaker identification model, a first threshold value and a second threshold value that can identify the same speaker are calculated. Also, at the time of speaker identification, the similarity between the calculated voice data to be identified and the registered voice data is corrected by subtracting the first threshold value or the second threshold value from the similarity. Then, by comparing the corrected similarity with a threshold value common to the first speaker identification model and the second speaker identification model, the speaker of the voice data to be identified can be identified with higher accuracy.

[0033] A speaker identification device according to another aspect of the present disclosure includes an identification target voice data acquisition unit that acquires voice data to be identified, a registered voice data acquisition unit that acquires registered voice data registered in advance, a first extraction unit that extracts a feature amount of the identification target voice data, a second extraction unit that extracts a feature amount of the registered voice data, a speaker identification model selection unit that selects a first speaker identification model learned by machine learning using male voice data to identify a male speaker when the gender of either the speaker of the identification target voice data or the speaker of the registered voice data is male, and selects a second speaker identification model learned by machine learning using female voice data to identify a female speaker when the gender of either the speaker of the identification target voice data or the speaker of the registered voice data is female, and a speaker identification unit that identifies the speaker of the identification target voice data by inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model.

[0034] According to this configuration, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the feature amounts of the voice data to be identified and the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the voice data to be identified is identified. Also, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the feature amounts of the voice data to be identified and the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the voice data to be identified is identified.

[0035] Therefore, even when the distribution of the feature amounts of the voice data differs depending on gender, the speaker of the voice data to be identified is identified by the first speaker identification model and the second speaker identification model specialized for each gender, so that the accuracy of identifying whether the speaker to be identified is a pre-registered speaker can be improved.

[0036] A speaker identification program according to another aspect of the present disclosure acquires voice data to be identified, acquires pre-registered registered voice data, extracts the feature amounts of the voice data to be identified, extracts the feature amounts of the registered voice data, and when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, selects the first speaker identification model that has been machine-learned using male voice data to identify male speakers, and when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, selects the second speaker identification model that has been machine-learned using female voice data to identify female speakers, and causes a computer to function so as to identify the speaker of the voice data to be identified by inputting the feature amounts of the voice data to be identified and the feature amounts of the registered voice data into either the selected first speaker identification model or the second speaker identification model.

[0037] According to this configuration, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the voice data to be identified is identified. Further, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the voice data to be identified is identified.

[0038] Therefore, even when the distribution of the feature amounts of the voice data differs depending on gender, the speaker of the voice data to be identified is identified by the first speaker identification model and the second speaker identification model specialized for each gender, so that the accuracy of identifying whether the speaker to be identified is a pre-registered speaker can be improved.

[0039] A gender identification model generation method according to another aspect of the present disclosure is such that a computer acquires a plurality of voice data to which gender labels indicating either male or female are assigned, and uses, as teacher data, the feature amounts of the first voice data and the second voice data among the plurality of voice data and the similarity of the gender labels of the first voice data and the second voice data, generates, by machine learning, a gender identification model in which the input is the feature amounts of two voice data and the output is the similarity of the two voice data.

[0040] According to this configuration, when the feature amount of the registered voice data or the voice data to be identified and the feature amount of the male voice data are input into the gender identification model generated by machine learning, a first similarity of the two voice data is output. Further, when the feature amount of the registered voice data or the voice data to be identified and the feature amount of the female voice data are input into the gender identification model, a second similarity of the two voice data is output. Then, by comparing the first similarity and the second similarity, the gender of the speaker of the registered voice data or the voice data to be identified can be easily estimated.

[0041] A speaker identification model generation method according to another aspect of the present disclosure includes a computer obtaining a plurality of male voice data with speaker identification labels for identifying male speakers, using, as teacher data, the feature amounts of first male voice data and second male voice data among the plurality of male voice data and the similarity of the speaker identification labels of the first male voice data and the second male voice data, generating, by machine learning, a first speaker identification model with the input being the feature amounts of two voice data and the output being the similarity of the two voice data, obtaining a plurality of female voice data with speaker identification labels for identifying female speakers, and using, as teacher data, the feature amounts of first female voice data and second female voice data among the plurality of female voice data and the similarity of the speaker identification labels of the first female voice data and the second female voice data, and generating, by machine learning, a second speaker identification model with the input being the feature amounts of two voice data and the output being the similarity of the two voice data.

[0042] According to this configuration, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the voice data to be identified is identified. Further, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the voice data to be identified is identified.

[0043] Therefore, even when the distribution of the feature amounts of the voice data differs depending on gender, the speaker of the voice data to be identified can be identified by the first speaker identification model specialized for males and the second speaker identification model specialized for females, so that the accuracy of identifying whether the speaker to be identified is a pre-registered speaker can be improved.

[0044] The embodiments of the present disclosure will be described with reference to the accompanying drawings below. Note that the following embodiments are examples that embody the present disclosure and do not limit the technical scope of the present disclosure.

[0045] (Embodiment 1) FIG. 1 is a diagram showing the configuration of a speaker identification system according to Embodiment 1 of the present disclosure.

[0046] The speaker identification system shown in FIG. 1 includes a microphone 1 and a speaker identification device 2. Note that the speaker identification device 2 may or may not include the microphone 1.

[0047] The microphone 1 picks up the voice uttered by the speaker, converts it into voice data, and outputs it to the speaker identification device 2. When registering voice data in advance, the microphone 1 outputs the registration target voice data uttered by the speaker to the speaker identification device 2. Also, when identifying the speaker, the microphone 1 outputs the identification target voice data uttered by the speaker to the speaker identification device 2.

[0048] The speaker identification device 2 includes a registration target voice data acquisition unit 201, a feature amount extraction unit 202, a gender identification model storage unit 203, a gender identification voice data storage unit 204, a gender identification unit 205, a registration unit 206, an identification target voice data acquisition unit 211, a registered voice data storage unit 212, a registered voice data acquisition unit 213, a feature amount extraction unit 214, a feature amount extraction unit 215, a speaker identification model storage unit 216, a model selection unit 217, a speaker identification unit 218, and an identification result output unit 219.

[0049] The registration target voice data acquisition unit 201, the feature amount extraction unit 202, the gender identification unit 205, the registration unit 206, the identification target voice data acquisition unit 211, the registered voice data acquisition unit 213, the feature amount extraction unit 214, the feature amount extraction unit 215, the model selection unit 217, the speaker identification unit 218, and the identification result output unit 219 are realized by a processor. The processor is composed of, for example, a CPU (Central Processing Unit).

[0050] The gender identification model storage unit 203, the voice data storage unit 204 for gender identification, the registered voice data storage unit 212, and the speaker identification model storage unit 216 are realized by a memory. The memory is composed of, for example, a ROM (Read Only Memory) or an EEPROM (Electrically Erasable Programmable Read Only Memory).

[0051] Note that the speaker identification device 2 may be, for example, a computer, a smartphone, a tablet computer, or a server.

[0052] The registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1.

[0053] The feature extraction unit 202 extracts the feature amount of the registration target voice data acquired by the registration target voice data acquisition unit 201. The feature amount is, for example, an i-vector. The i-vector is a feature amount of a low-dimensional vector extracted from voice data by using factor analysis on a GMM (Gaussian Mixture Model) supervector. Since the method for extracting the i-vector is a prior art, a detailed description thereof is omitted. Also, the feature amount is not limited to the i-vector, and may be other feature amounts such as an x-vector, for example.

[0054] The gender identification model storage unit 203 stores in advance a gender identification model that has been machine-learned using voice data of men and women in order to identify the gender of a speaker. Note that the method for generating the gender identification model will be described later.

[0055] The voice data storage unit 204 for gender identification stores in advance the feature amounts of the voice data for gender identification used to identify the gender of the speaker of the voice data to be registered. The voice data for gender identification includes a plurality of voice data of men and a plurality of voice data of women. Although the voice data storage unit 204 for gender identification stores in advance the feature amounts of the voice data for gender identification, the present disclosure is not particularly limited thereto, and the voice data for gender identification may be stored in advance. In this case, the speaker identification device 2 includes a feature amount extraction unit that extracts the feature amounts of the voice data for gender identification.

[0056] The gender identification unit 205 identifies the gender of the speaker of the voice data to be registered using the feature amounts of the voice data to be registered extracted by the feature amount extraction unit 202. The gender identification unit 205 acquires from the gender identification model storage unit 203 a gender identification model that has been machine-learned using voice data of men and women to identify the gender of the speaker. The gender identification unit 205 identifies the gender of the speaker of the voice data to be registered by inputting the feature amounts of the voice data to be registered into the gender identification model.

[0057] The gender identification unit 205 inputs the feature amounts of the voice data to be registered and the feature amounts of each of the plurality of voice data of men stored in advance in the voice data storage unit 204 for gender identification into the gender identification model, thereby obtaining from the gender identification model the similarity between the voice data to be registered and each of the plurality of voice data of men. Then, the gender identification unit 205 calculates the average of the obtained plurality of similarities as the average male similarity.

[0058] Also, the gender identification unit 205 inputs the feature amounts of the voice data to be registered and the feature amounts of each of the plurality of voice data of women stored in advance in the voice data storage unit 204 for gender identification into the gender identification model, thereby obtaining from the gender identification model the similarity between the voice data to be registered and each of the plurality of voice data of women. Then, the gender identification unit 205 calculates the average of the obtained plurality of similarities as the average female similarity.

[0059] When the average male similarity is higher than the average female similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as male. On the other hand, when the average male similarity is lower than the average female similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as female. Note that when the average male similarity is the same as the average female similarity, the gender identification unit 205 may identify the gender of the speaker of the voice data to be registered as male, or may identify the gender of the speaker of the voice data to be registered as female.

[0060] The registration unit 206 registers the voice data to be registered with the gender information associated by the gender identification unit 205 as registered voice data. The registration unit 206 registers the registered voice data in the registered voice data storage unit 212.

[0061] Note that the speaker identification device 2 may further include an input reception unit that receives an input of information regarding the speaker of the voice data to be registered. Then, the registration unit 206 may register the registered voice data in the registered voice data storage unit 212 in association with the information regarding the speaker. The information regarding the speaker is, for example, the name of the speaker.

[0062] The identification target voice data acquisition unit 211 acquires the identification target voice data output from the microphone 1.

[0063] The registered voice data storage unit 212 stores the registered voice data associated with the gender information. The registered voice data storage unit 212 stores a plurality of registered voice data.

[0064] The registered voice data acquisition unit 213 acquires the registered voice data registered in the registered voice data storage unit 212.

[0065] The feature amount extraction unit 214 extracts the feature amount of the identification target voice data acquired by the identification target voice data acquisition unit 211. The feature amount is, for example, an i-vector.

[0066] The feature extraction unit 215 extracts the features of the registered voice data acquired by the registered voice data acquisition unit 213. The feature is, for example, an i-vector.

[0067] The speaker identification model storage unit 216 stores in advance a first speaker identification model that is machine-learned using male voice data to identify male speakers and a second speaker identification model that is machine-learned using female voice data to identify female speakers. The speaker identification model storage unit 216 stores in advance the first speaker identification model and the second speaker identification model generated by a speaker identification model generation device 4 described later. Note that the generation method of the first speaker identification model and the second speaker identification model will be described later.

[0068] When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the model selection unit 217 selects the first speaker identification model that is machine-learned using male voice data to identify male speakers. Also, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the model selection unit 217 selects the second speaker identification model that is machine-learned using female voice data to identify female speakers.

[0069] In the first embodiment, when the gender of the speaker of the registered voice data is male, the model selection unit 217 selects the first speaker identification model, and when the gender of the speaker of the registered voice data is female, the model selection unit 217 selects the second speaker identification model. The gender is associated in advance with the registered voice data. Therefore, when the gender associated with the registered voice data is male, the model selection unit 217 selects the first speaker identification model, and when the gender associated with the registered voice data is female, the model selection unit 217 selects the second speaker identification model.

[0070] The speaker identification unit 218 identifies the speaker of the voice data to be identified by inputting the features of the voice data to be identified and the features of the registered voice data into either the first speaker identification model or the second speaker identification model selected by the model selection unit 217.

[0071] The speaker identification unit 218 includes a similarity calculation unit 231 and a similarity determination unit 232.

[0072] The similarity calculation unit 231 inputs the feature amount of the voice data to be identified and the feature amounts of a plurality of registered voice data respectively into either the selected first speaker identification model or the second speaker identification model, and obtains the similarity between the voice data to be identified and each of the plurality of registered voice data from either the first speaker identification model or the second speaker identification model.

[0073] The similarity determination unit 232 identifies the speaker of the registered voice data with the highest obtained similarity as the speaker of the voice data to be identified.

[0074] Note that the similarity determination unit 232 may determine whether the highest similarity is greater than a threshold value. Even when there is no registered voice data of the same speaker as the speaker of the voice data to be identified in the registered voice data storage unit 212, the similarity between the voice data to be identified and each registered voice data is calculated. Therefore, even for the registered voice data with the highest similarity, the speaker of the registered voice data is not necessarily the same as the speaker of the voice data to be identified. Thus, by determining whether the highest similarity is greater than the threshold value, the speaker can be reliably identified.

[0075] The identification result output unit 219 outputs the identification result by the speaker identification unit 218. The identification result output unit 219 is, for example, a display or a speaker. When the speaker of the voice data to be identified is identified, the identification result output unit 219 outputs a message indicating that the speaker of the voice data to be identified is a pre-registered speaker to the display or the speaker. On the other hand, when the speaker of the voice data to be identified is not identified, the identification result output unit 219 outputs a message indicating that the speaker of the voice data to be identified is not a pre-registered speaker to the display or the speaker. The identification result output unit 219 may output the identification result by the speaker identification unit 218 to another device other than the speaker identification device 2.

[0076] Next, a gender identification model generation device according to Embodiment 1 of the present disclosure will be described.

[0077] FIG. 2 is a diagram showing the configuration of a gender identification model generation device according to Embodiment 1 of the present disclosure.

[0078] The gender identification model generation device 3 shown in FIG. 2 includes a voice data storage unit 301 for gender identification, a voice data acquisition unit 302 for gender identification, a feature extraction unit 303, a gender identification model generation unit 304, and a gender identification model storage unit 305.

[0079] The voice data acquisition unit 302 for gender identification, the feature extraction unit 303, and the gender identification model generation unit 304 are realized by a processor. The voice data storage unit 301 for gender identification and the gender identification model storage unit 305 are realized by a memory.

[0080] The voice data storage unit 301 for gender identification pre-stores a plurality of voice data with gender labels indicating either male or female. The voice data storage unit 301 for gender identification stores a plurality of different voice data for each of a plurality of speakers.

[0081] The voice data acquisition unit 302 for gender identification acquires a plurality of voice data with gender labels indicating either male or female from the voice data storage unit 301 for gender identification. In this Embodiment 1, the voice data acquisition unit 302 for gender identification acquires a plurality of voice data from the voice data storage unit 301 for gender identification, but the present disclosure is not particularly limited thereto, and a plurality of voice data may be acquired (received) from an external device via a network.

[0082] The feature extraction unit 303 extracts feature amounts of a plurality of voice data acquired by the voice data acquisition unit 302 for gender identification. The feature amount is, for example, an i-vector.

[0083] The gender identification model generation unit 304 uses, as teacher data, the feature amounts of the first voice data and the second voice data among a plurality of voice data and the similarity of the gender labels of the first voice data and the second voice data, takes the inputs as the feature amounts of two voice data, and generates, by machine learning, a gender identification model with the outputs being the similarity of the two voice data. For example, if the gender label of the first voice data and the gender label of the second voice data are the same, the gender identification model outputs the highest similarity, and if the gender label of the first voice data and the gender label of the second voice data are different, the gender identification model is machine-learned to output the lowest similarity.

[0084] As the gender identification model, a model based on Probabilistic Linear Discriminant Analysis (PLDA) is used. The PLDA model automatically selects feature amounts effective for speaker identification from 400-dimensional i-vector feature amounts and calculates the log-likelihood ratio as the similarity.

[0085] Note that examples of machine learning include supervised learning in which the relationship between inputs and outputs is learned using teacher data with labels (output information) assigned to the input information, unsupervised learning in which the data structure is constructed from only unlabeled inputs, semi-supervised learning that handles both labeled and unlabeled data, and reinforcement learning in which actions to maximize rewards are learned through trial and error. Further, specific methods of machine learning include neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, or support vector machines (SVM). In the machine learning of the gender identification model, any of the specific examples listed above may be used.

[0086] The gender identification model storage unit 305 stores the gender identification model generated by the gender identification model generation unit 304.

[0087] Incidentally, the gender identification model generation device 3 may transmit the gender identification model stored in the gender identification model storage unit 305 to the speaker identification device 2. The speaker identification device 2 may store the received gender identification model in the gender identification model storage unit 203. Also, at the time of manufacturing the speaker identification device 2, the gender identification model generated by the gender identification model generation device 3 may be stored in the speaker identification device 2.

[0088] In the gender identification model generation device 3 of the first embodiment, gender labels indicating either male or female are attached to the plurality of voice data, but the present disclosure is not particularly limited thereto, and identification information for identifying a speaker may be attached as a label. In this case, the gender identification model generation unit 304 uses, as teacher data, the feature amounts of the first voice data and the second voice data among the plurality of voice data and the similarity of the identification information of the first voice data and the second voice data, and generates, by machine learning, a gender identification model in which the input is the feature amounts of the two voice data and the output is the similarity of the two voice data. For example, the gender identification model is machine-learned such that the highest similarity is output if the identification information of the first voice data and the identification information of the second voice data are the same, and the lowest similarity is output if the identification information of the first voice data and the identification information of the second voice data are different.

[0089] Subsequently, the speaker identification model generation device in the first embodiment of the present disclosure will be described.

[0090] FIG. 3 is a diagram showing the configuration of the speaker identification model generation device in the first embodiment of the present disclosure.

[0091] The speaker identification model generation device 4 shown in FIG. 3 includes a male voice data storage unit 401, a male voice data acquisition unit 402, a feature amount extraction unit 403, a first speaker identification model generation unit 404, a first speaker identification model storage unit 405, a female voice data storage unit 411, a female voice data acquisition unit 412, a feature amount extraction unit 413, a second speaker identification model generation unit 414, and a second speaker identification model storage unit 415.

[0092] The male voice data acquisition unit 402, the feature amount extraction unit 403, the first speaker identification model generation unit 404, the female voice data acquisition unit 412, the feature amount extraction unit 413, and the second speaker identification model generation unit 414 are realized by a processor. The male voice data storage unit 401, the first speaker identification model storage unit 405, the female voice data storage unit 411, and the second speaker identification model storage unit 415 are realized by a memory.

[0093] The male voice data storage unit 401 stores a plurality of male voice data with speaker identification labels attached thereto for identifying speakers who are male. The male voice data storage unit 401 stores a plurality of different male voice data for each of a plurality of speakers.

[0094] The male voice data acquisition unit 402 acquires a plurality of male voice data with speaker identification labels attached thereto for identifying speakers who are male from the male voice data storage unit 401. In the first embodiment, the male voice data acquisition unit 402 acquires a plurality of male voice data from the male voice data storage unit 401. However, the present disclosure is not particularly limited thereto, and a plurality of male voice data may be acquired (received) from an external device via a network.

[0095] The feature amount extraction unit 403 extracts the feature amounts of the plurality of male voice data acquired by the male voice data acquisition unit 402. The feature amount is, for example, an i-vector.

[0096] The first speaker identification model generation unit 404 uses, as teacher data, the feature amounts of the first male voice data and the second male voice data among the plurality of male voice data and the similarity of the speaker identification labels of the first male voice data and the second male voice data, takes the input as the feature amounts of two voice data, and generates, by machine learning, a first speaker identification model in which the output is the similarity of the two voice data. For example, the first speaker identification model is machine-learned such that if the speaker identification labels of the first male voice data and the second male voice data are the same, the highest similarity is output, and if the speaker identification labels of the first male voice data and the second male voice data are different, the lowest similarity is output.

[0097] As the first speaker identification model, a model based on PLDA is used. The PLDA model automatically selects effective features for speaker identification from 400-dimensional i-vectors (features), and calculates the log-likelihood ratio as the similarity.

[0098] Note that as machine learning, for example, supervised learning that learns the relationship between input and output using training data with labels (output information) attached to the input information, unsupervised learning that constructs the structure of data from only unlabeled input, semi-supervised learning that handles both labeled and unlabeled data, reinforcement learning that learns actions to maximize rewards through trial and error, etc. can be mentioned. Also, as specific methods of machine learning, there are neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, or support vector machines (SVM), etc. In the machine learning of the first speaker identification model, any of the specific examples listed above may be used.

[0099] The first speaker identification model storage unit 405 stores the first speaker identification model generated by the first speaker identification model generation unit 404.

[0100] The female voice data storage unit 411 stores a plurality of female voice data with speaker identification labels for identifying speakers who are female. The female voice data storage unit 411 stores a plurality of different female voice data for each of the plurality of speakers.

[0101] The female voice data acquisition unit 412 acquires a plurality of female voice data with speaker identification labels for identifying speakers who are female from the female voice data storage unit 411. Note that in the first embodiment, the female voice data acquisition unit 412 acquires a plurality of female voice data from the female voice data storage unit 411, but the present disclosure is not particularly limited thereto, and a plurality of female voice data may be acquired (received) from an external device via a network.

[0102] The feature extraction unit 413 extracts the features of a plurality of female voice data acquired by the female voice data acquisition unit 412. The feature is, for example, an i-vector.

[0103] The second speaker identification model generation unit 414 uses, as teacher data, the features of the first female voice data and the second female voice data among the plurality of female voice data and the similarity of the speaker identification labels of the first female voice data and the second female voice data, takes the input as the features of two voice data, and generates, by machine learning, a second speaker identification model whose output is the similarity of the two voice data. For example, if the speaker identification label of the first female voice data and the speaker identification label of the second female voice data are the same, the second speaker identification model outputs the highest similarity by machine learning, and if the speaker identification label of the first female voice data and the speaker identification label of the second female voice data are different, the second speaker identification model outputs the lowest similarity.

[0104] As the second speaker identification model, a model based on PLDA is used. The PLDA model automatically selects features effective for speaker identification from 400-dimensional i-vectors (features) and calculates the log-likelihood ratio as the similarity.

[0105] Note that examples of machine learning include supervised learning in which the relationship between input and output is learned using teacher data with a label (output information) attached to the input information, unsupervised learning in which the data structure is constructed from only the input without a label, semi-supervised learning that handles both with and without labels, and reinforcement learning in which actions that maximize rewards are learned through trial and error. Also, specific methods of machine learning include neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, or support vector machines (SVMs). Any of the specific examples listed above may be used in the machine learning of the second speaker identification model.

[0106] The second speaker identification model storage unit 415 stores the second speaker identification model generated by the second speaker identification model generation unit 414.

[0107] Note that the speaker identification model generation device 4 may transmit the first speaker identification model stored in the first speaker identification model storage unit 405 and the second speaker identification model stored in the second speaker identification model storage unit 415 to the speaker identification device 2. The speaker identification device 2 may store the received first speaker identification model and second speaker identification model in the speaker identification model storage unit 216. Also, at the time of manufacturing the speaker identification device 2, the first speaker identification model and the second speaker identification model generated by the speaker identification model generation device 4 may be stored in the speaker identification device 2.

[0108] Subsequently, the operations of the registration process and the speaker identification process of the speaker identification device 2 in the first embodiment will be described.

[0109] FIG. 4 is a flowchart for explaining the operation of the registration process of the speaker identification device in the first embodiment.

[0110] First, in step S1, the registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1. A speaker who wishes to register the voice data of his or her own utterance speaks a predetermined sentence toward the microphone 1. At this time, it is preferable that the sentence of the registration target voice data is longer than the sentence of the identification target voice data. By acquiring registration target voice data with a relatively large number of characters, the accuracy of speaker identification can be improved. Also, the speaker identification device 2 may present a plurality of predetermined sentences to the registration target speaker. In this case, the registration target speaker speaks the plurality of presented sentences.

[0111] Next, in step S2, the feature extraction unit 202 extracts the features of the registration target voice data acquired by the registration target voice data acquisition unit 201.

[0112] Next, in step S3, the gender identification unit 205 performs gender identification processing to identify the gender of the speaker of the registration target voice data using the feature amount of the registration target voice data extracted by the feature extraction unit 202. The gender identification processing will be described later.

[0113] Next, in step S4, the registration unit 206 stores the registration target voice data associated with the gender information identified by the gender identification unit 205 in the registered voice data storage unit 212 as registered voice data.

[0114] FIG. 5 is a flowchart for explaining the operation of the gender identification processing in step S3 of FIG. 4.

[0115] First, in step S11, the gender identification unit 205 acquires a gender identification model from the gender identification model storage unit 203. The gender identification model is a gender identification model that is machine-learned using, as teacher data, the feature amounts of the first voice data and the second voice data among a plurality of voice data and the similarity between the gender labels of the first voice data and the second voice data. Note that the gender identification model may be a gender identification model that is machine-learned using, as teacher data, the feature amounts of the first voice data and the second voice data among a plurality of voice data and the similarity between the identification information of the first voice data and the second voice data.

[0116] Next, in step S12, the gender identification unit 205 acquires the feature amount of male voice data from the gender identification voice data storage unit 204.

[0117] Next, in step S13, the gender identification unit 205 calculates the similarity between the registration target voice data and the male voice data by inputting the feature amount of the registration target voice data and the feature amount of the male voice data into the gender identification model.

[0118] Next, in step S14, the gender discrimination unit 205 determines whether the similarity between the voice data to be registered and all male voice data has been calculated. Here, if it is determined that the similarity between the voice data to be registered and all male voice data has not been calculated (NO in step S14), the process returns to step S12. Then, the gender discrimination unit 205 acquires, from the gender discrimination voice data storage unit 204, the feature amount of the male voice data for which the similarity has not been calculated, from among the feature amounts of the plurality of male voice data.

[0119] On the other hand, if it is determined that the similarity between the voice data to be registered and all male voice data has been calculated (YES in step S14), in step S15, the gender discrimination unit 205 calculates the average of the calculated plurality of similarities as the average male similarity.

[0120] Next, in step S16, the gender discrimination unit 205 acquires the feature amount of the female voice data from the gender discrimination voice data storage unit 204.

[0121] Next, in step S17, the gender discrimination unit 205 inputs the feature amount of the voice data to be registered and the feature amount of the female voice data into the gender discrimination model, thereby calculating the similarity between the voice data to be registered and the female voice data.

[0122] Next, in step S18, the gender discrimination unit 205 determines whether the similarity between the voice data to be registered and all female voice data has been calculated. Here, if it is determined that the similarity between the voice data to be registered and all female voice data has not been calculated (NO in step S18), the process returns to step S16. Then, the gender discrimination unit 205 acquires, from the gender discrimination voice data storage unit 204, the feature amount of the female voice data for which the similarity has not been calculated, from among the feature amounts of the plurality of female voice data.

[0123] On the other hand, when it is determined that the similarity between the voice data to be registered and all the female voice data has been calculated (YES in step S18), in step S19, the gender discrimination unit 205 calculates the average of the calculated multiple similarities as the average female similarity.

[0124] Next, in step S20, the gender discrimination unit 205 outputs the gender with the higher value of the average male similarity and the average female similarity to the registration unit 206 as the discrimination result.

[0125] FIG. 6 is a first flowchart for explaining the operation of the speaker discrimination process of the speaker discrimination device in the first embodiment, and FIG. 7 is a second flowchart for explaining the operation of the speaker discrimination process of the speaker discrimination device in the first embodiment.

[0126] First, in step S31, the discrimination target voice data acquisition unit 211 acquires the discrimination target voice data output from the microphone 1. The discrimination target speaker speaks toward the microphone 1. The microphone 1 collects the voice spoken by the discrimination target speaker and outputs the discrimination target voice data.

[0127] Next, in step S32, the feature quantity extraction unit 214 extracts the feature quantity of the discrimination target voice data acquired by the discrimination target voice data acquisition unit 211.

[0128] Next, in step S33, the registered voice data acquisition unit 213 acquires the registered voice data from the registered voice data storage unit 212. At this time, the registered voice data acquisition unit 213 acquires one registered voice data from among the plurality of registered voice data registered in the registered voice data storage unit 212.

[0129] Next, in step S34, the feature quantity extraction unit 215 extracts the feature quantity of the registered voice data acquired by the registered voice data acquisition unit 213.

[0130] Next, in step S35, the model selection unit 217 acquires the gender associated with the registered voice data acquired by the registered voice data acquisition unit 213.

[0131] Next, in step S36, the model selection unit 217 determines whether the acquired gender is male. Here, if it is determined that the acquired gender is male (YES in step S36), in step S37, the model selection unit 217 selects the first speaker identification model. The model selection unit 217 acquires the selected first speaker identification model from the speaker identification model storage unit 216, and outputs the acquired first speaker identification model to the similarity calculation unit 231.

[0132] On the other hand, if it is determined that the acquired gender is not male, that is, if it is determined that the acquired gender is female (NO in step S36), in step S38, the model selection unit 217 selects the second speaker identification model. The model selection unit 217 acquires the selected second speaker identification model from the speaker identification model storage unit 216, and outputs the acquired second speaker identification model to the similarity calculation unit 231.

[0133] Next, in step S39, the similarity calculation unit 231 calculates the similarity between the identification target voice data and the registered voice data by inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model.

[0134] Next, in step S40, the similarity calculation unit 231 determines whether the similarity between the identification target voice data and all the registered voice data stored in the registered voice data storage unit 212 has been calculated. Here, if it is determined that the similarity between the identification target voice data and all the registered voice data has not been calculated (NO in step S40), the process returns to step S33. Then, the registered voice data acquisition unit 213 acquires the registered voice data for which the similarity has not been calculated from among the plurality of registered voice data stored in the registered voice data storage unit 212.

[0135] On the other hand, when it is determined that the similarity between the voice data to be identified and all the registered voice data has been calculated (YES in step S40), in step S41, the similarity determination unit 232 determines whether the highest similarity is greater than the threshold value.

[0136] Here, when it is determined that the highest similarity is greater than the threshold value (YES in step S41), in step S42, the similarity determination unit 232 identifies the speaker of the registered voice data with the highest similarity as the speaker of the voice data to be identified.

[0137] On the other hand, when it is determined that the highest similarity is less than or equal to the threshold value (NO in step S41), in step S43, the similarity determination unit 232 determines that there is no registered voice data among the plurality of registered voice data whose speaker is the same as that of the voice data to be identified.

[0138] Next, in step S44, the identification result output unit 219 outputs the identification result by the speaker identification unit 218. When the speaker of the voice data to be identified is identified, the identification result output unit 219 outputs a message indicating that the speaker of the voice data to be identified is a pre-registered speaker. On the other hand, when the speaker of the voice data to be identified is not identified, the identification result output unit 219 outputs a message indicating that the speaker of the voice data to be identified is not a pre-registered speaker.

[0139] In this way, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the voice data to be identified is identified. Also, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the voice data to be identified is identified.

[0140] Therefore, even when the distribution of the feature amounts of the voice data differs depending on the gender, the speaker of the voice data to be identified is identified by the first speaker identification model and the second speaker identification model specialized for each gender. Thus, it is possible to improve the accuracy of identifying whether or not the speaker to be identified is a speaker registered in advance.

[0141] Note that, in the first embodiment, the model selection unit 217 selects one of the first speaker identification model and the second speaker identification model based on the gender of the speaker of the registered voice data. However, the present disclosure is not particularly limited thereto. The model selection unit 217 may select one of the first speaker identification model and the second speaker identification model based on the gender of the speaker of the voice data to be identified. In this case, the speaker identification device 2 includes a gender identification unit that identifies the gender of the speaker of the voice data to be identified, a gender identification model storage unit that stores in advance a gender identification model that has been machine-learned using voice data of men and women in order to identify the gender of the speaker, and a gender identification voice data storage unit that stores in advance the feature amounts of the gender identification voice data used to identify the gender of the speaker of the voice data to be identified. The gender identification unit, the gender identification model storage unit, and the gender identification voice data storage unit have the same configuration as the above-described gender identification unit 205, gender identification model storage unit 203, and gender identification voice data storage unit 204. Further, when the gender of the speaker of the voice data to be identified is identified, the gender identification unit 205, the gender identification model storage unit 203, and the gender identification voice data storage unit 204 become unnecessary.

[0142] Next, the operation of the gender identification model generation process of the gender identification model generation device 3 in the first embodiment will be described.

[0143] FIG. 8 is a flowchart for explaining the operation of the gender identification model generation process of the gender identification model generation device in the first embodiment.

[0144] First, in step S51, the gender identification voice data acquisition unit 302 acquires a plurality of voice data to which a gender label indicating either male or female is attached from the gender identification voice data storage unit 301.

[0145] Next, in step S52, the feature amount extraction unit 303 extracts the feature amounts of the plurality of voice data acquired by the gender identification voice data acquisition unit 302.

[0146] Next, in step S53, the gender identification model generation unit 304 acquires, as teacher data, the feature amounts of the first voice data and the second voice data among the plurality of voice data and the similarity of the gender labels of the first voice data and the second voice data.

[0147] Next, in step S54, the gender identification model generation unit 304 performs machine learning on a gender identification model that uses the acquired teacher data, takes the feature amounts of two voice data as inputs, and takes the similarity of the two voice data as outputs.

[0148] Next, in step S55, the gender identification model generation unit 304 determines whether or not machine learning has been performed on the gender identification model using all combinations of the plurality of voice data. Here, if it is determined that machine learning has not been performed on the gender identification model using all combinations of the voice data (NO in step S55), the process returns to step S53. Then, the gender identification model generation unit 304 acquires, as teacher data, the feature amounts of the first voice data and the second voice data of the combination not used in the machine learning among the plurality of voice data and the similarity of the gender labels of the first voice data and the second voice data.

[0149] On the other hand, if it is determined that machine learning has been performed on the gender identification model using all combinations of the voice data (YES in step S55), in step S56, the gender identification model generation unit 304 stores the gender identification model generated by the machine learning in the gender identification model storage unit 305.

[0150] In this way, by inputting the feature amounts of the registered voice data or the voice data to be identified and the feature amounts of the male voice data into the gender identification model generated by machine learning, the first similarity of the two voice data is output. Further, by inputting the feature amounts of the registered voice data or the voice data to be identified and the feature amounts of the female voice data into the gender identification model, the second similarity of the two voice data is output. Then, by comparing the first similarity and the second similarity, the gender of the speaker of the registered voice data or the voice data to be identified can be easily estimated.

[0151] Subsequently, the operation of the speaker identification model generation process of the speaker identification model generation device 4 in the first embodiment will be described.

[0152] FIG. 9 is a first flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in the first embodiment, and FIG. 10 is a second flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in the first embodiment.

[0153] First, in step S61, the male voice data acquisition unit 402 acquires a plurality of male voice data with speaker identification labels attached for identifying a speaker who is male from the male voice data storage unit 401.

[0154] Next, in step S62, the feature amount extraction unit 403 extracts the feature amounts of the plurality of male voice data acquired by the male voice data acquisition unit 402.

[0155] Next, in step S63, the first speaker identification model generation unit 404 acquires, as teacher data, the feature amounts of the first male voice data and the second male voice data among the plurality of male voice data and the similarities of the speaker identification labels of the first male voice data and the second male voice data.

[0156] Next, in step S64, the first speaker identification model generation unit 404 performs machine learning on a first speaker identification model that uses the acquired teacher data, with the input being the feature amounts of two pieces of voice data and the output being the similarity of the two pieces of voice data.

[0157] Next, in step S65, the first speaker identification model generation unit 404 determines whether or not it has performed machine learning on the first speaker identification model using all combinations of the plurality of male voice data. Here, if it is determined that machine learning has not been performed on the first speaker identification model using all combinations of the male voice data (NO in step S65), the process returns to step S63. Then, the first speaker identification model generation unit 404 acquires, as teacher data, the feature amounts of the first male voice data and the second male voice data of the combination not used in machine learning among the plurality of male voice data, and the similarity of the speaker identification labels of the first male voice data and the second male voice data.

[0158] On the other hand, if it is determined that machine learning has been performed on the first speaker identification model using all combinations of the male voice data (YES in step S65), in step S66, the first speaker identification model generation unit 404 stores the first speaker identification model generated by machine learning in the first speaker identification model storage unit 405.

[0159] Next, in step S67, the female voice data acquisition unit 412 acquires, from the female voice data storage unit 411, a plurality of pieces of female voice data to which speaker identification labels for identifying a speaker who is female are assigned.

[0160] Next, in step S68, the feature amount extraction unit 413 extracts the feature amounts of the plurality of pieces of female voice data acquired by the female voice data acquisition unit 412.

[0161] Next, in step S69, the second speaker identification model generation unit 414 acquires, as teacher data, the feature amounts of the first female voice data and the second female voice data among the plurality of female voice data, and the similarity of the speaker identification labels of the first female voice data and the second female voice data.

[0162] Next, in step S70, the second speaker identification model generation unit 414 uses the acquired teacher data to perform machine learning on a second speaker identification model in which the input is the feature amount of each of the two voice data and the output is the similarity of the two voice data.

[0163] Next, in step S71, the second speaker identification model generation unit 414 determines whether or not the second speaker identification model has been machine-learned using combinations of all of the plurality of female voice data. Here, if it is determined that the second speaker identification model has not been machine-learned using combinations of all of the female voice data (NO in step S71), the process returns to step S69. Then, the second speaker identification model generation unit 414 acquires, as teacher data, the feature amounts of the first female voice data and the second female voice data of the combination not used in the machine learning among the plurality of female voice data, and the similarity of the speaker identification labels of the first female voice data and the second female voice data.

[0164] On the other hand, if it is determined that the second speaker identification model has been machine-learned using combinations of all of the female voice data (YES in step S71), in step S72, the second speaker identification model generation unit 414 stores the second speaker identification model generated by the machine learning in the second speaker identification model storage unit 415.

[0165] Thus, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the first speaker identification model generated for males, whereby the speaker of the voice data to be identified is identified. Further, when the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the second speaker identification model generated for females, whereby the speaker of the voice data to be identified is identified.

[0166] Therefore, even when the distribution of the feature amounts of the voice data differs depending on gender, the speaker of the voice data to be identified can be identified by the first speaker identification model specialized for males and the second speaker identification model specialized for females, so that the accuracy of identifying whether the speaker to be identified is a pre-registered speaker can be improved.

[0167] Subsequently, the evaluation of the speaker identification performance of the speaker identification device 2 in the first embodiment will be described.

[0168] FIG. 11 is a diagram showing the speaker identification performance evaluation results of a conventional speaker identification device and the speaker identification performance evaluation results of the speaker identification device of the first embodiment.

[0169] The performance evaluation results shown in FIG. 11 represent the results of identifying the speakers of the voice data provided by the SRE19 progress dataset and the SRE19 Evaluation dataset using a conventional speaker identification device and the speaker identification device 2 of the first embodiment.

[0170] SRE19 is a speaker identification competition hosted by the National Institute of Standards and Technology (NIST) of the United States. The SRE19 progress dataset and the SRE19 Evaluation dataset are both datasets provided by SRE19.

[0171] A conventional speaker identification device identified the speaker of voice data using a speaker identification model generated without distinguishing between male and female speakers.

[0172] Also, the speaker identification device 2 according to Embodiment 1 identified the speaker of voice data using a first speaker identification model for males and a second speaker identification model for females.

[0173] The evaluation results were represented by the EER (Equal Error Rate) (%) commonly used for evaluating speaker identification, the minC (minimum Cost) which is the cost defined by NIST, and the actC (actual Cost) which is the cost defined by NIST. Note that for EER, minC, and actC, the smaller the value, the higher the performance.

[0174] As shown in FIG. 11, the EER, minC, and actC of the speaker identification device 2 according to Embodiment 1 are all smaller than the EER, minC, and actC of the conventional speaker identification device. From this, it can be seen that the speaker identification performance of the speaker identification device 2 according to Embodiment 1 is higher than the speaker identification performance of the conventional speaker identification device.

[0175] Next, the speaker identification device 2 in Modification 1 of Embodiment 1 will be described.

[0176] In the speaker identification device 2 in Modification 1 of Embodiment 1, the gender identification process of the gender identification unit 205 is different from the above.

[0177] The gender identification unit 205 in Modification 1 of Embodiment 1 inputs the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data stored in advance in the gender identification voice data storage unit 204 into the gender identification model, thereby obtaining the similarity between the voice data to be registered and each of the plurality of male voice data from the gender identification model. Then, the gender identification unit 205 calculates the maximum value among the obtained plurality of similarities as the maximum male similarity.

[0178] Further, the gender identification unit 205 inputs the feature amount of the voice data to be registered and the feature amounts of a plurality of voice data of women pre-stored in the voice data storage unit 204 for gender identification into the gender identification model, thereby obtaining the similarity between the voice data to be registered and each of the plurality of voice data of women from the gender identification model. Then, the gender identification unit 205 calculates the maximum value among the obtained plurality of similarities as the maximum female similarity.

[0179] If the maximum male similarity is higher than the maximum female similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as male. On the other hand, if the maximum male similarity is lower than the maximum female similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as female. Note that when the maximum male similarity is the same as the maximum female similarity, the gender identification unit 205 may identify the gender of the speaker of the voice data to be registered as male or may identify the gender of the speaker of the voice data to be registered as female.

[0180] Subsequently, the operation of the gender identification process in Modification 1 of the first embodiment will be described.

[0181] FIG. 12 is a flowchart for explaining the operation of the gender identification process in Modification 1 of the first embodiment. The gender identification process in Modification 1 of the first embodiment is another operation example of the gender identification process in step S3 of FIG. 4.

[0182] Since the processes of steps S81 to S84 are the same as the processes of steps S11 to S14 shown in FIG. 5, the description thereof will be omitted.

[0183] When it is determined that the similarities between the voice data to be registered and all the male voice data have been calculated (YES in step S84), in step S85, the gender identification unit 205 calculates the maximum value among the calculated plurality of similarities as the maximum male similarity.

[0184] Since the processes of steps S86 to S88 are the same as the processes of steps S16 to S18 shown in FIG. 5, the description thereof will be omitted.

[0185] When it is determined that the similarity between the voice data to be registered and all the female voice data has been calculated (YES in step S88), in step S89, the gender discrimination unit 205 calculates the maximum value among the calculated multiple similarities as the maximum female similarity.

[0186] Next, in step S90, the gender discrimination unit 205 outputs the gender with the higher value between the maximum male similarity and the maximum female similarity to the registration unit 206 as the discrimination result.

[0187] Subsequently, the speaker discrimination device 2 in the second modification example of the first embodiment will be described.

[0188] In the speaker discrimination device 2 in the second modification example of the first embodiment, the gender discrimination process of the gender discrimination unit 205 is different from the above.

[0189] The gender discrimination unit 205 in the second modification example of the first embodiment calculates the average feature amounts of a plurality of male voice data stored in advance. Then, the gender discrimination unit 205 inputs the feature amount of the voice data to be registered and the average feature amounts of the plurality of male voice data into the gender discrimination model, and obtains the first similarity between the voice data to be registered and the group of male voice data from the gender discrimination model.

[0190] Also, the gender discrimination unit 205 calculates the average feature amounts of a plurality of female voice data stored in advance. Then, the gender discrimination unit 205 inputs the feature amount of the voice data to be registered and the average feature amounts of the plurality of female voice data into the gender discrimination model, and obtains the second similarity between the voice data to be registered and the group of female voice data from the gender discrimination model.

[0191] When the first similarity is higher than the second similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as male. On the other hand, when the first similarity is lower than the second similarity, the gender identification unit 205 identifies the gender of the speaker of the voice data to be registered as female. Note that when the first similarity is the same as the second similarity, the gender identification unit 205 may identify the gender of the speaker of the voice data to be registered as male, or may identify the gender of the speaker of the voice data to be registered as female.

[0192] Next, the operation of the gender identification process in Modification 2 of Embodiment 1 will be described.

[0193] FIG. 13 is a flowchart for explaining the operation of the gender identification process in Modification 2 of Embodiment 1. The gender identification process in Modification 2 of Embodiment 1 is another operation example of the gender identification process in step S3 of FIG. 4.

[0194] First, in step S101, the gender identification unit 205 acquires a gender identification model from the gender identification model storage unit 203. The gender identification model is a gender identification model that is machine-learned using, as teacher data, the feature amounts of the first voice data and the second voice data among a plurality of voice data, and the similarity of the gender labels of the first voice data and the second voice data. Note that the gender identification model may be a gender identification model that is machine-learned using, as teacher data, the feature amounts of the first voice data and the second voice data among a plurality of voice data, and the similarity of the identification information of the first voice data and the second voice data.

[0195] Next, in step S102, the gender identification unit 205 acquires the feature amounts of a plurality of male voice data from the voice data storage unit 204 for gender identification.

[0196] Next, in step S103, the gender identification unit 205 calculates the average feature amount of the acquired plurality of male voice data.

[0197] Next, in step S104, the gender discrimination unit 205 calculates a first similarity between the voice data to be registered and the male voice data group by inputting the feature amount of the voice data to be registered and the average feature amount of a plurality of male voice data into the gender discrimination model.

[0198] Next, in step S105, the gender discrimination unit 205 acquires the feature amounts of a plurality of female voice data from the voice data storage unit 204 for gender discrimination.

[0199] Next, in step S106, the gender discrimination unit 205 calculates the average feature amount of the acquired plurality of female voice data.

[0200] Next, in step S107, the gender discrimination unit 205 calculates a second similarity between the voice data to be registered and the female voice data group by inputting the feature amount of the voice data to be registered and the average feature amount of a plurality of female voice data into the gender discrimination model.

[0201] Next, in step S108, the gender discrimination unit 205 outputs, to the registration unit 206, as the discrimination result, the gender with the higher of the first similarity and the second similarity.

[0202] In the gender discrimination process in the first embodiment, the similarity between the registered voice data and the male voice data and the similarity between the registered voice data and the female voice data are calculated, and the gender of the registered voice data is determined by comparing the two calculated similarities. However, the present disclosure is not particularly limited to this. For example, the gender discrimination model generation unit 304 may use, as teacher data, the feature amount of one voice data among a plurality of voice data and the gender label of the one voice data, use the input as the feature amount of the voice data, and generate a gender discrimination model whose output is either male or female by machine learning. In this case, the gender discrimination model is, for example, a deep neural network model, and the machine learning is, for example, deep learning.

[0203] Further, the gender identification model may calculate the probability of being male and the probability of being female for the input voice data. In this case, the gender identification unit 205 may output, as the identification result, the gender with the higher probability among the probability of being male and the probability of being female.

[0204] Further, the speaker identification device 2 may include an input reception unit that receives an input of the gender of the speaker of the voice data to be registered when the voice data to be registered is acquired. In this case, the registration unit 206 may store the voice data to be registered acquired by the voice data to be registered acquisition unit 201 in the registered voice data storage unit 212 in association with the gender received by the input reception unit. Thereby, the feature amount extraction unit 202, the gender identification model storage unit 203, the voice data storage unit 204 for gender identification, and the gender identification unit 205 become unnecessary, the configuration of the speaker identification device 2 can be simplified, and the load on the registration process of the speaker identification device 2 can be reduced.

[0205] (Embodiment 2) In the above-described Embodiment 1, the similarity between each of all the registered voice data stored in the registered voice data storage unit 212 and the voice data to be identified is calculated, and the speaker of the registered voice data with the highest similarity is identified as the speaker of the voice data to be identified. On the other hand, in Embodiment 2, identification information of the speaker of the voice data to be identified is input, and one registered voice data that is preliminarily associated with the identification information is acquired from among a plurality of registered voice data stored in the registered voice data storage unit 212. Then, the similarity between the one registered voice data and the voice data to be identified is calculated, and when the similarity is higher than the threshold value, the speaker of the registered voice data is identified as the speaker of the voice data to be identified.

[0206] FIG. 14 is a diagram showing the configuration of a speaker identification system according to Embodiment 2 of the present disclosure.

[0207] The speaker identification system shown in FIG. 14 includes a microphone 1 and a speaker identification device 21. Note that the speaker identification device 21 may or may not include the microphone 1.

[0208] In addition, in the second embodiment, the same components as those in the first embodiment are denoted by the same reference numerals, and the description thereof is omitted.

[0209] The speaker identification device 21 includes a registration target voice data acquisition unit 201, a feature amount extraction unit 202, a gender identification model storage unit 203, a voice data storage unit for gender identification 204, a gender identification unit 205, a registration unit 2061, an identification target voice data acquisition unit 211, a registered voice data storage unit 2121, a registered voice data acquisition unit 2131, a feature amount extraction unit 214, a feature amount extraction unit 215, a speaker identification model storage unit 216, a model selection unit 217, a speaker identification unit 2181, an identification result output unit 219, an input reception unit 221, and an identification information acquisition unit 222.

[0210] The input reception unit 221 is an input device such as a keyboard, a mouse, and a touch panel. At the time of registering voice data, the input reception unit 221 receives an input by the speaker of the identification information for identifying the speaker who registers the voice data. Also, at the time of identifying voice data, the input reception unit 221 receives an input by the speaker of the identification information for identifying the speaker of the voice data to be identified. Note that the input reception unit 221 may be a card reader or an RFID (Radio Frequency IDentification) reader or the like. In this case, the speaker causes the card reader to read a card on which the identification information is recorded, or causes the RFID reader to read an RFID tag on which the identification information is recorded.

[0211] The identification information acquisition unit 222 acquires the identification information received by the input reception unit 221. At the time of registering voice data, the identification information acquisition unit 222 acquires the identification information for identifying the speaker of the registration target voice data, and outputs the acquired identification information to the registration unit 2061. Also, at the time of identifying voice data, the identification information acquisition unit 222 acquires the identification information for identifying the speaker of the identification target voice data, and outputs the acquired identification information to the registered voice data acquisition unit 2131.

[0212] The registration unit 2061 registers the registration target voice data in which the gender information identified by the gender identification unit 205 and the identification information obtained by the identification information acquisition unit 222 are associated as registered voice data. The registration unit 2061 registers the registered voice data in the registered voice data storage unit 2121.

[0213] The registered voice data storage unit 2121 stores the registered voice data in which the gender information and the identification information are associated. The registered voice data storage unit 2121 stores a plurality of registered voice data. The plurality of registered voice data are associated with the identification information for identifying the speaker of each of the plurality of registered voice data.

[0214] The registered voice data acquisition unit 2131 acquires, from among the plurality of registered voice data registered in the registered voice data storage unit 2121, the registered voice data associated with the identification information that matches the identification information acquired by the identification information acquisition unit 222.

[0215] The speaker identification unit 2181 includes a similarity calculation unit 2311 and a similarity determination unit 2321.

[0216] The similarity calculation unit 2311 obtains the similarity between the identification target voice data and the registered voice data from either the first speaker identification model or the second speaker identification model by inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model.

[0217] When the obtained similarity is higher than the threshold value, the similarity determination unit 2321 identifies the speaker of the registered voice data as the speaker of the identification target voice data.

[0218] Subsequently, the operations of the registration process and the speaker identification process of the speaker identification device 21 in the second embodiment will be described.

[0219] FIG. 15 is a flowchart for explaining the operation of the registration process of the speaker identification device in the second embodiment.

[0220] First, in step S121, the identification information acquisition unit 222 acquires the identification information of the speaker received by the input reception unit 221. The input reception unit 221 receives the input of the identification information by the speaker for identifying the speaker who registers the voice data, and outputs the received identification information to the identification information acquisition unit 222. The identification information acquisition unit 222 outputs the identification information for identifying the speaker of the voice data to be registered to the registration unit 2061.

[0221] Next, in step S122, the voice data acquisition unit 201 to be registered acquires the voice data to be registered output from the microphone 1. The speaker who wishes to register the voice data of his or her own speech for which the identification information has been input speaks a predetermined sentence toward the microphone 1.

[0222] The processes of step S123 and step S124 are the same as the processes of step S2 and step S3 shown in FIG. 4, and thus the description thereof is omitted.

[0223] Next, in step S125, the registration unit 2061 stores, as registered voice data, the voice data to be registered in which the gender information identified by the gender identification unit 205 and the identification information acquired by the identification information acquisition unit 222 are associated, in the registered voice data storage unit 2121. Thereby, the registered voice data storage unit 2121 stores the registered voice data in which the gender information and the identification information are associated.

[0224] FIG. 16 is a first flowchart for explaining the operation of the speaker identification process of the speaker identification device in the second embodiment, and FIG. 17 is a second flowchart for explaining the operation of the speaker identification process of the speaker identification device in the second embodiment.

[0225] First, in step S131, the identification information acquisition unit 222 acquires the identification information of the speaker received by the input reception unit 221. The input reception unit 221 receives the input of the identification information by the speaker for identifying the speaker of the voice data to be identified, and outputs the received identification information to the identification information acquisition unit 222. The identification information acquisition unit 222 outputs the identification information for identifying the speaker of the identification target voice data to the registered voice data acquisition unit 2131.

[0226] Next, in step S132, the registered voice data acquisition unit 2131 determines whether there is identification information in the registered voice data storage unit 2121 that matches the identification information acquired by the identification information acquisition unit 222. Here, if it is determined that there is no identification information in the registered voice data storage unit 2121 that matches the acquired identification information (NO in step S132), the speaker identification process ends. Note that if it is determined that there is no identification information in the registered voice data storage unit 2121 that matches the acquired identification information, the identification result output unit 219 may output notification information for notifying the speaker that the input identification information is not registered. Also, if it is determined that there is no identification information in the registered voice data storage unit 2121 that matches the acquired identification information, the identification result output unit 219 may output notification information for prompting the speaker to register the voice data.

[0227] On the other hand, if it is determined that there is identification information in the registered voice data storage unit 2121 that matches the acquired identification information (YES in step S132), in step S133, the identification target voice data acquisition unit 211 acquires the identification target voice data output from the microphone 1.

[0228] Next, in step S134, the feature amount extraction unit 214 extracts the feature amount of the identification target voice data acquired by the identification target voice data acquisition unit 211.

[0229] Next, in step S135, the registered voice data acquisition unit 2131 acquires the registered voice data associated with the identification information obtained by the identification information acquisition unit 222 from among the plurality of registered voice data registered in the registered voice data storage unit 2121.

[0230] The processes of steps S136 to S141 are the same as the processes of steps S34 to S39 shown in FIG. 6, and thus the description thereof is omitted.

[0231] Next, in step S142, the similarity determination unit 2321 determines whether the similarity calculated by the similarity calculation unit 2311 is greater than the threshold value.

[0232] Here, when it is determined that the similarity calculated by the similarity calculation unit 2311 is greater than the threshold value (YES in step S142), in step S143, the similarity determination unit 2321 identifies the speaker of the registered voice data as the speaker of the identification target voice data.

[0233] On the other hand, when it is determined that the similarity calculated by the similarity calculation unit 2311 is less than or equal to the threshold value (NO in step S142), in step S144, the similarity determination unit 2321 determines that the speaker of the identification target voice data is not the speaker of the registered voice data.

[0234] Next, in step S145, the identification result output unit 219 outputs the identification result by the speaker identification unit 2181. When the speaker of the identification target voice data is identified, the identification result output unit 219 outputs a message indicating that the speaker of the identification target voice data is a pre-registered speaker. On the other hand, when the speaker of the identification target voice data is not identified, the identification result output unit 219 outputs a message indicating that the speaker of the identification target voice data is not a pre-registered speaker.

[0235] Thus, in the second embodiment, only the similarity between the registered voice data associated with the identification information and the voice data to be identified is calculated. Therefore, compared with the first embodiment in which a plurality of similarities between each of the plurality of registered voice data and the voice data to be identified are calculated, the processing load of similarity calculation can be reduced in the second embodiment.

[0236] (Embodiment 3) In the above-described first and second embodiments, either the first speaker identification model or the second speaker identification model is selected according to the gender of the speaker of the registered voice data. However, when two different first speaker identification models and second speaker identification models are used, the output value ranges of the first speaker identification model and the second speaker identification model may be different. Therefore, in the third embodiment, at the time of registering voice data, for each of the first speaker identification model and the second speaker identification model, a first threshold value and a second threshold value capable of identifying the same speaker are calculated, and the calculated first threshold value and second threshold value are stored. Also, at the time of identifying voice data, the similarity is corrected by subtracting the first threshold value or the second threshold value from the calculated similarity between the voice data to be identified and the registered voice data. Then, the corrected similarity is compared with a third threshold value common to the first speaker identification model and the second speaker identification model, whereby the speaker of the voice data to be identified is identified.

[0237] First, the speaker identification model generation device in the third embodiment of the present disclosure will be described.

[0238] FIG. 18 is a diagram showing the configuration of the speaker identification model generation device in the third embodiment of the present disclosure.

[0239] The speaker identification model generation device 41 shown in FIG. 18 includes a male voice data storage unit 401, a male voice data acquisition unit 402, a feature amount extraction unit 403, a first speaker identification model generation unit 404, a first speaker identification model storage unit 405, a first speaker identification unit 406, a first threshold calculation unit 407, a threshold storage unit 408, a female voice data storage unit 411, a female voice data acquisition unit 412, a feature amount extraction unit 413, a second speaker identification model generation unit 414, a second speaker identification model storage unit 415, a second speaker identification unit 416, and a second threshold calculation unit 417.

[0240] In addition, in the third embodiment, the same components as those in the first and second embodiments are denoted by the same reference numerals, and the description thereof is omitted.

[0241] The first speaker identification unit 406 inputs all combinations of the feature amounts of two pieces of voice data out of a plurality of male voice data into the first speaker identification model, and obtains the similarity of each of the plurality of combinations of the two pieces of voice data from the first speaker identification model.

[0242] The first threshold calculation unit 407 calculates a first threshold that can distinguish the similarity of two pieces of voice data of the same speaker from the similarity of two pieces of voice data of different speakers. The first threshold calculation unit 407 calculates the first threshold by performing regression analysis on the plurality of similarities calculated by the first speaker identification unit 406.

[0243] The second speaker identification unit 416 inputs all combinations of the feature amounts of two pieces of voice data out of a plurality of female voice data into the second speaker identification model, and obtains the similarity of each of the plurality of combinations of the two pieces of voice data from the second speaker identification model.

[0244] The second threshold calculation unit 417 calculates a second threshold that can distinguish the similarity of two pieces of voice data of the same speaker from the similarity of two pieces of voice data of different speakers. The second threshold calculation unit 417 calculates the second threshold by performing regression analysis on the plurality of similarities calculated by the second speaker identification unit 416.

[0245] The threshold memory unit 408 stores the first threshold calculated by the first threshold calculation unit 407 and the second threshold calculated by the second threshold calculation unit 417.

[0246] Subsequently, the speaker identification system in Embodiment 3 of the present disclosure will be described.

[0247] FIG. 19 is a diagram showing the configuration of the speaker identification system in Embodiment 3 of the present disclosure.

[0248] The speaker identification system shown in FIG. 19 includes a microphone 1 and a speaker identification device 22. Note that the speaker identification device 22 may or may not include the microphone 1.

[0249] In Embodiment 3, the same components as those in Embodiment 1 and Embodiment 2 are denoted by the same reference numerals, and the description thereof is omitted.

[0250] The speaker identification device 22 includes a registration target voice data acquisition unit 201, a feature amount extraction unit 202, a gender identification model memory unit 203, a voice data memory unit 204 for gender identification, a gender identification unit 205, a registration unit 2061, an identification target voice data acquisition unit 211, a registered voice data memory unit 2121, a registered voice data acquisition unit 2131, a feature amount extraction unit 214, a feature amount extraction unit 215, a speaker identification model memory unit 216, a model selection unit 217, a speaker identification unit 2182, an identification result output unit 219, an input reception unit 221, an identification information acquisition unit 222, and a threshold memory unit 223.

[0251] The speaker identification unit 2182 includes a similarity calculation unit 2311, a similarity correction unit 233, and a similarity determination unit 2322.

[0252] When the similarity correction unit 233 obtains the similarity from the first speaker identification model, it subtracts the first threshold value from the obtained similarity. Also, when the similarity correction unit 233 obtains the similarity from the second speaker identification model, it subtracts the second threshold value from the obtained similarity. When the similarity is calculated using the first speaker identification model by the similarity calculation unit 2311, the similarity correction unit 233 reads the first threshold value from the threshold storage unit 223 and subtracts the first threshold value from the calculated similarity. Also, when the similarity is calculated using the second speaker identification model by the similarity calculation unit 2311, the similarity correction unit 233 reads the second threshold value from the threshold storage unit 223 and subtracts the second threshold value from the calculated similarity.

[0253] The threshold storage unit 223 stores in advance a first threshold value for correcting the similarity calculated using the first speaker identification model and a second threshold value for correcting the similarity calculated using the second speaker identification model. The threshold storage unit 223 stores in advance the first threshold value and the second threshold value generated by the speaker identification model generation device 41.

[0254] Note that the speaker identification model generation device 41 may transmit the first threshold value and the second threshold value stored in the threshold storage unit 408 to the speaker identification device 22. The speaker identification device 22 may store the received first threshold value and second threshold value in the threshold storage unit 223. Also, at the time of manufacturing the speaker identification device 22, the first threshold value and the second threshold value generated by the speaker identification model generation device 41 may be stored in the threshold storage unit 223.

[0255] When the similarity from which the first threshold value or the second threshold value has been subtracted by the similarity correction unit 233 is higher than the third threshold value, the similarity determination unit 2322 identifies the speaker of the registered voice data as the speaker of the identification target voice data.

[0256] Subsequently, the operation of the speaker identification model generation process of the speaker identification model generation device 41 in the third embodiment will be described.

[0257] FIG. 20 is a first flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 3. FIG. 21 is a second flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 3. FIG. 22 is a third flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 3.

[0258] The processes of steps S151 to S156 are the same as the processes of steps S61 to S66 shown in FIG. 9, so the description thereof is omitted.

[0259] Next, in step S157, the first speaker identification unit 406 acquires the first speaker identification model from the first speaker identification model storage unit 405.

[0260] Next, in step S158, the first speaker identification unit 406 acquires the feature amounts of two male voice data from among the feature amounts of a plurality of male voice data extracted by the feature amount extraction unit 403.

[0261] Next, in step S159, the first speaker identification unit 406 calculates the similarity of the two male voice data by inputting the acquired feature amounts of the two male voice data into the first speaker identification model. Note that the two male voice data are either two male voice data uttered by one speaker or two male voice data uttered by two speakers. At this time, the similarity when the two male voice data are male voice data uttered by one speaker is higher than the similarity when the two male voice data are male voice data uttered by two speakers.

[0262] Next, in step S160, the first speaker identification unit 406 determines whether the similarity of all combinations of male voice data has been calculated. Here, if it is determined that the similarity of all combinations of male voice data has not been calculated (NO in step S160), the process returns to step S158. Then, the first speaker identification unit 406 acquires, from the feature extraction unit 403, the feature amounts of two male voice data for which the similarity has not been calculated, from among the feature amounts of a plurality of male voice data.

[0263] On the other hand, if it is determined that the similarity of all combinations of male voice data has been calculated (YES in step S160), in step S161, the first threshold calculation unit 407 calculates a first threshold that can distinguish the similarity between two male voice data of the same speaker and the similarity between two male voice data of different speakers, by performing a regression analysis on the plurality of similarities calculated by the first speaker identification unit 406.

[0264] Next, in step S162, the first threshold calculation unit 407 stores the calculated first threshold in the threshold storage unit 408.

[0265] The processes of steps S163 to S168 are the same as the processes of steps S67 to S72 shown in FIGS. 9 and 10, and thus the description thereof is omitted.

[0266] Next, in step S169, the second speaker identification unit 416 acquires the second speaker identification model from the second speaker identification model storage unit 415.

[0267] Next, in step S170, the second speaker identification unit 416 acquires the feature amounts of two female voice data from among the feature amounts of a plurality of female voice data extracted by the feature extraction unit 413.

[0268] Next, in step S171, the second speaker identification unit 416 calculates the similarity of the two pieces of female voice data by inputting the feature amounts of the acquired two pieces of female voice data into the second speaker identification model. Note that the two pieces of female voice data are either two pieces of female voice data spoken by one speaker or two pieces of female voice data spoken by two speakers. At this time, the similarity when the two pieces of female voice data are female voice data spoken by one speaker is higher than the similarity when the two pieces of female voice data are female voice data spoken by two speakers.

[0269] Next, in step S172, the second speaker identification unit 416 determines whether the similarities of all combinations of the female voice data have been calculated. Here, if it is determined that the similarities of all combinations of the female voice data have not been calculated (NO in step S172), the process returns to step S170. Then, the second speaker identification unit 416 acquires, from the feature extraction unit 413, the feature amounts of two pieces of female voice data for which the similarity has not been calculated, from among the feature amounts of the plurality of pieces of female voice data.

[0270] On the other hand, if it is determined that the similarities of all combinations of the female voice data have been calculated (YES in step S172), in step S173, the second threshold calculation unit 417 calculates a second threshold that can distinguish the similarity between two pieces of female voice data of the same speaker and the similarity between two pieces of female voice data of different speakers, by performing regression analysis on the plurality of similarities calculated by the second speaker identification unit 416.

[0271] Next, in step S174, the second threshold calculation unit 417 stores the calculated second threshold in the threshold storage unit 408.

[0272] Subsequently, the operation of the speaker identification process of the speaker identification device 22 in the third embodiment will be described.

[0273] FIG. 23 is a first flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 3, and FIG. 24 is a second flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 3.

[0274] The processes of steps S181 to S191 are the same as the processes of steps S131 to S141 shown in FIGS. 16 and 17, and thus the description thereof is omitted.

[0275] Next, in step S192, the similarity correction unit 233 acquires the first threshold value or the second threshold value from the threshold value storage unit 223. At this time, when the first speaker identification model is selected by the model selection unit 217, the similarity correction unit 233 acquires the first threshold value from the threshold value storage unit 223. Also, when the second speaker identification model is selected by the model selection unit 217, the similarity correction unit 233 acquires the second threshold value from the threshold value storage unit 223.

[0276] Next, in step S193, the similarity correction unit 233 corrects the similarity calculated by the similarity calculation unit 2311 using the acquired first threshold value or second threshold value. At this time, the similarity correction unit 233 subtracts the first threshold value or the second threshold value from the similarity calculated by the similarity calculation unit 2311.

[0277] Next, in step S194, the similarity determination unit 2322 determines whether the similarity corrected by the similarity correction unit 233 is greater than the third threshold value. Note that the third threshold value is, for example, 0. When the corrected similarity is greater than 0, the similarity determination unit 2322 determines that the voice data to be identified matches the registered voice data registered in advance, and when the corrected similarity is 0 or less, the similarity determination unit 2322 determines that the voice data to be identified does not match the registered voice data registered in advance.

[0278] Here, when it is determined that the similarity corrected by the similarity correction unit 233 is greater than the third threshold (YES in step S194), in step S195, the similarity determination unit 2322 identifies the speaker of the registered voice data as the speaker of the voice data to be identified.

[0279] On the other hand, when it is determined that the similarity corrected by the similarity correction unit 233 is less than or equal to the third threshold (NO in step S194), in step S196, the similarity determination unit 2322 determines that the speaker of the voice data to be identified is not the speaker of the registered voice data.

[0280] The process of step S197 is the same as the process of step S145 shown in FIG. 17, so the description thereof is omitted.

[0281] When two different first speaker identification models and second speaker identification models are used, the output value ranges of the first speaker identification model and the second speaker identification model may be different. Therefore, in the third embodiment, at the time of registration, first thresholds and second thresholds capable of identifying the same speaker are calculated for the first speaker identification model and the second speaker identification model, respectively. Also, at the time of speaker identification, the similarity is corrected by subtracting the first threshold or the second threshold from the calculated similarity between the voice data to be identified and the registered voice data. Then, by comparing the corrected similarity with a third threshold common to the first speaker identification model and the second speaker identification model, the speaker of the voice data to be identified can be identified with higher accuracy.

[0282] (Embodiment 4) Since the longer the speaking time, the more information there is in the speech, it is easier to identify the speaker, and the similarity between the registered voice data of the speakers themselves and the voice data to be identified tends to be high. On the other hand, since the shorter the speaking time, the less information there is in the speech, it is difficult to identify the speaker, and there is a possibility that the similarity is low even between the registered voice data of the speakers themselves and the voice data to be identified. Therefore, when identifying the speaker of the voice data to be identified with a short speaking time using a speaker identification model trained with voice data having a long speaking time, the accuracy of speaker identification may decrease.

[0283] Therefore, in the speaker identification method according to Embodiment 4, a computer acquires voice data to be identified, acquires registered voice data registered in advance, extracts feature amounts of the voice data to be identified, extracts feature amounts of the registered voice data, and when at least one of the speaking times of the speaker of the voice data to be identified and the speaker of the registered voice data is equal to or longer than a predetermined time, selects a third speaker identification model trained using voice data with a speaking time equal to or longer than the predetermined time to identify the speaker with a speaking time equal to or longer than the predetermined time, and when at least one of the speaking times of the speaker of the voice data to be identified and the speaker of the registered voice data is shorter than the predetermined time, selects a fourth speaker identification model trained using voice data with a speaking time shorter than the predetermined time to identify the speaker with a speaking time shorter than the predetermined time, and inputs the feature amounts of the voice data to be identified and the feature amounts of the registered voice data into either the selected third speaker identification model or the fourth speaker identification model to identify the speaker of the voice data to be identified.

[0284] FIG. 25 is a diagram showing the configuration of a speaker identification system according to Embodiment 4 of the present disclosure.

[0285] The speaker identification system shown in FIG. 25 includes a microphone 1 and a speaker identification device 24. Note that the speaker identification device 24 may or may not include the microphone 1.

[0286] In addition, in Embodiment 4, the same components as those in Embodiment 1 are denoted by the same reference numerals, and the description thereof is omitted.

[0287] The speaker identification device 24 includes a registration target voice data acquisition unit 201, a speech time measurement unit 207, a registration unit 2064, an identification target voice data acquisition unit 211, a registered voice data storage unit 2124, a registered voice data acquisition unit 213, a feature quantity extraction unit 214, a feature quantity extraction unit 215, a speaker identification model storage unit 2164, a model selection unit 2174, a speaker identification unit 2184, and an identification result output unit 219.

[0288] The speech time measurement unit 207 measures the speech time of the registration target voice data acquired by the registration target voice data acquisition unit 201. Note that the speech time is the time from the time when the acquisition of the registration target voice data is started by the registration target voice data acquisition unit 201 to the time when the acquisition of the registration target voice data is completed.

[0289] The registration unit 2064 registers the registration target voice data associated with the speech time information indicating the speech time measured by the speech time measurement unit 207 as registered voice data. The registration unit 2064 registers the registered voice data in the registered voice data storage unit 2124.

[0290] Note that the speaker identification device 24 may further include an input reception unit that receives an input of information regarding the speaker of the registration target voice data. Then, the registration unit 2064 may register the registered voice data in the registered voice data storage unit 2124 in association with the information regarding the speaker. The information regarding the speaker is, for example, the name of the speaker.

[0291] The registered voice data storage unit 2124 stores the registered voice data associated with the speech time information. The registered voice data storage unit 2124 stores a plurality of registered voice data.

[0292] The speaker identification model storage unit 2164 stores in advance a third speaker identification model that is machine-learned using voice data with a speaking time of a predetermined time or longer to identify a speaker with a speaking time of a predetermined time or longer, and a fourth speaker identification model that is machine-learned using voice data with a speaking time shorter than the predetermined time to identify a speaker with a speaking time shorter than the predetermined time. The speaker identification model storage unit 2164 stores in advance the third speaker identification model and the fourth speaker identification model generated by a speaker identification model generation device 44 described later. Note that the generation methods of the third speaker identification model and the fourth speaker identification model will be described later.

[0293] When at least one of the speaking times of the speaker of the voice data to be identified and the speaker of the registered voice data is a predetermined time or longer, the model selection unit 2174 selects a third speaker identification model that is machine-learned using voice data with a speaking time of a predetermined time or longer to identify a speaker with a speaking time of a predetermined time or longer. Also, when at least one of the speaking times of the speaker of the voice data to be identified and the speaker of the registered voice data is shorter than the predetermined time, the model selection unit 2174 selects a fourth speaker identification model that is machine-learned using voice data with a speaking time shorter than the predetermined time to identify a speaker with a speaking time shorter than the predetermined time.

[0294] In the fourth embodiment, when the speaking time of the speaker of the registered voice data is a predetermined time or longer, the model selection unit 2174 selects the third speaker identification model, and when the speaking time of the speaker of the registered voice data is shorter than the predetermined time, the model selection unit 2174 selects the fourth speaker identification model. The speaking time is associated with the registered voice data in advance. Therefore, when the speaking time associated with the registered voice data is a predetermined time or longer, the model selection unit 2174 selects the third speaker identification model, and when the speaking time associated with the registered voice data is shorter than the predetermined time, the model selection unit 2174 selects the fourth speaker identification model. Note that the predetermined time is, for example, 60 seconds.

[0295] The speaker identification unit 2184 identifies the speaker of the voice data to be identified by inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the third speaker identification model or the fourth speaker identification model selected by the model selection unit 2174.

[0296] The speaker identification unit 2184 includes a similarity calculation unit 2314 and a similarity determination unit 232.

[0297] The similarity calculation unit 2314 obtains the similarity between the voice data to be identified and each of the plurality of registered voice data from either the third speaker identification model or the fourth speaker identification model by inputting the feature amount of the voice data to be identified and the feature amount of each of the plurality of registered voice data into either the selected third speaker identification model or the fourth speaker identification model.

[0298] Subsequently, a speaker identification model generation device according to Embodiment 4 of the present disclosure will be described.

[0299] FIG. 26 is a diagram showing the configuration of a speaker identification model generation device according to Embodiment 4 of the present disclosure.

[0300] The speaker identification model generation device 44 shown in FIG. 26 includes a long-term voice data storage unit 421, a long-term voice data acquisition unit 422, a feature amount extraction unit 423, a third speaker identification model generation unit 424, a third speaker identification model storage unit 425, a short-term voice data storage unit 431, a short-term voice data acquisition unit 432, a feature amount extraction unit 433, a fourth speaker identification model generation unit 434, and a fourth speaker identification model storage unit 435.

[0301] The long-term voice data acquisition unit 422, the feature amount extraction unit 423, the third speaker identification model generation unit 424, the short-term voice data acquisition unit 432, the feature amount extraction unit 433, and the fourth speaker identification model generation unit 434 are realized by a processor. The long-term voice data storage unit 421, the third speaker identification model storage unit 425, the short-term voice data storage unit 431, and the fourth speaker identification model storage unit 435 are realized by a memory.

[0302] The long-duration voice data storage unit 421 stores a plurality of long-duration voice data with a speaker identification label for identifying a speaker and having a speech time of a predetermined time or longer. The long-duration voice data is voice data having a speech time of a predetermined time or longer. The long-duration voice data storage unit 421 stores a plurality of different long-duration voice data for each of a plurality of speakers.

[0303] The long-duration voice data acquisition unit 422 acquires a plurality of long-duration voice data with a speaker identification label for identifying a speaker from the long-duration voice data storage unit 421. In the fourth embodiment, the long-duration voice data acquisition unit 422 acquires a plurality of long-duration voice data from the long-duration voice data storage unit 421. However, the present disclosure is not particularly limited thereto, and a plurality of long-duration voice data may be acquired (received) from an external device via a network.

[0304] The feature amount extraction unit 423 extracts feature amounts of a plurality of long-duration voice data acquired by the long-duration voice data acquisition unit 422. The feature amount is, for example, an i-vector.

[0305] The third speaker identification model generation unit 424 uses, as teacher data, the feature amounts of the first long-duration voice data and the second long-duration voice data among the plurality of long-duration voice data and the similarity of the speaker identification labels of the first long-duration voice data and the second long-duration voice data, uses the input as the feature amounts of two voice data, and generates, by machine learning, a third speaker identification model in which the output is the similarity of the two voice data. For example, the third speaker identification model is machine-learned such that if the speaker identification label of the first long-duration voice data is the same as the speaker identification label of the second long-duration voice data, the highest similarity is output, and if the speaker identification label of the first long-duration voice data is different from the speaker identification label of the second long-duration voice data, the lowest similarity is output.

[0306] As the third speaker identification model, a model based on PLDA is used. The PLDA model automatically selects feature amounts effective for speaker identification from 400-dimensional i-vectors (feature amounts) and calculates the log-likelihood ratio as the similarity.

[0307] Note that as machine learning, for example, there are supervised learning that learns the relationship between input and output using training data with labels (output information) assigned to the input information, unsupervised learning that constructs the structure of data from only the input without labels, semi-supervised learning that deals with both with and without labels, reinforcement learning that learns actions to maximize rewards through trial and error, and the like. Also, as specific methods of machine learning, there are neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, or support vector machines (SVM), etc. In the machine learning of the third speaker identification model, any of the specific examples listed above may be used.

[0308] The third speaker identification model storage unit 425 stores the third speaker identification model generated by the third speaker identification model generation unit 424.

[0309] The short-time voice data storage unit 431 stores a plurality of short-time voice data with a speaker identification label for identifying a speaker and having a speech time shorter than a predetermined time. The short-time voice data is voice data having a speech time shorter than a predetermined time. The short-time voice data storage unit 431 stores a plurality of short-time voice data different from each other for each of a plurality of speakers.

[0310] The short-time voice data acquisition unit 432 acquires a plurality of short-time voice data with a speaker identification label for identifying a speaker from the short-time voice data storage unit 431. Note that in the fourth embodiment, the short-time voice data acquisition unit 432 acquires a plurality of short-time voice data from the short-time voice data storage unit 431, but the present disclosure is not particularly limited thereto, and a plurality of short-time voice data may be acquired (received) from an external device via a network.

[0311] The feature extraction unit 433 extracts features of the plurality of short-time voice data acquired by the short-time voice data acquisition unit 432. The feature is, for example, an i-vector.

[0312] The fourth speaker identification model generation unit 434 uses, as teacher data, the feature amounts of the first short-time voice data and the second short-time voice data among a plurality of short-time voice data, and the similarity of the speaker identification labels of the first short-time voice data and the second short-time voice data, takes the inputs as the feature amounts of two voice data, and generates, by machine learning, a fourth speaker identification model in which the output is the similarity of the two voice data. For example, if the speaker identification label of the first short-time voice data and the speaker identification label of the second short-time voice data are the same, the highest similarity is output by the fourth speaker identification model, and if the speaker identification label of the first short-time voice data and the speaker identification label of the second short-time voice data are different, the lowest similarity is output by machine learning.

[0313] As the fourth speaker identification model, a model based on PLDA is used. The PLDA model automatically selects feature amounts effective for speaker identification from 400-dimensional i-vectors (feature amounts), and calculates the log likelihood ratio as the similarity.

[0314] Note that examples of machine learning include supervised learning in which the relationship between inputs and outputs is learned using teacher data with labels (output information) given to the input information, unsupervised learning in which the data structure is constructed from only unlabeled inputs, semi-supervised learning that handles both labeled and unlabeled data, and reinforcement learning in which actions that maximize rewards are learned through trial and error. Also, specific methods of machine learning include neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, or support vector machines (SVMs). In the machine learning of the fourth speaker identification model, any of the specific examples listed above may be used.

[0315] The fourth speaker identification model storage unit 435 stores the fourth speaker identification model generated by the fourth speaker identification model generation unit 434.

[0316] Note that the speaker identification model generation device 44 may transmit the third speaker identification model stored in the third speaker identification model storage unit 425 and the fourth speaker identification model stored in the fourth speaker identification model storage unit 435 to the speaker identification device 24. The speaker identification device 24 may store the received third speaker identification model and fourth speaker identification model in the speaker identification model storage unit 2164. Further, at the time of manufacturing the speaker identification device 24, the third speaker identification model and the fourth speaker identification model generated by the speaker identification model generation device 44 may be stored in the speaker identification device 24.

[0317] Subsequently, the operations of the registration process and the speaker identification process of the speaker identification device 24 in the fourth embodiment will be described.

[0318] FIG. 27 is a flowchart for explaining the operation of the registration process of the speaker identification device in the fourth embodiment.

[0319] First, in step S201, the registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1. A speaker who wishes to register the voice data of his or her own utterance speaks a predetermined sentence toward the microphone 1. At this time, the predetermined sentence is either a sentence whose utterance time is equal to or longer than a predetermined time or a sentence whose utterance time is shorter than the predetermined time. The speaker identification device 24 may present a plurality of predetermined sentences to the registration target speaker. In this case, the registration target speaker speaks the plurality of presented sentences.

[0320] Next, in step S202, the utterance time measurement unit 207 measures the utterance time of the registration target voice data acquired by the registration target voice data acquisition unit 201.

[0321] Next, in step S203, the registration unit 2064 stores the registration target voice data associated with the utterance time information indicating the utterance time measured by the utterance time measurement unit 207 in the registered voice data storage unit 2124 as registered voice data.

[0322] FIG. 28 is a first flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 4, and FIG. 29 is a second flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 4.

[0323] The processes of steps S211 to S214 are the same as the processes of steps S31 to S34 shown in FIG. 6, so the description thereof is omitted.

[0324] Next, in step S215, the model selection unit 2174 acquires the utterance time associated with the registered voice data acquired by the registered voice data acquisition unit 213.

[0325] Next, in step S216, the model selection unit 2174 determines whether the acquired utterance time is equal to or longer than a predetermined time. Here, if it is determined that the acquired utterance time is equal to or longer than the predetermined time (YES in step S216), in step S217, the model selection unit 2174 selects the third speaker identification model. The model selection unit 2174 acquires the selected third speaker identification model from the speaker identification model storage unit 2164, and outputs the acquired third speaker identification model to the similarity calculation unit 2314.

[0326] On the other hand, if it is determined that the acquired utterance time is not equal to or longer than the predetermined time, that is, if it is determined that the acquired utterance time is shorter than the predetermined time (NO in step S216), in step S218, the model selection unit 2174 selects the fourth speaker identification model. The model selection unit 2174 acquires the selected fourth speaker identification model from the speaker identification model storage unit 2164, and outputs the acquired fourth speaker identification model to the similarity calculation unit 2314.

[0327] Next, in step S219, the similarity calculation unit 2314 calculates the similarity between the identification target voice data and the registered voice data by inputting the feature amount of the identification target voice data and the feature amount of the registered voice data into either the selected third speaker identification model or the fourth speaker identification model.

[0328] Next, in step S220, the similarity calculation unit 2314 determines whether the similarity between the voice data to be identified and all the registered voice data stored in the registered voice data storage unit 2124 has been calculated. Here, if it is determined that the similarity between the voice data to be identified and all the registered voice data has not been calculated (NO in step S220), the process returns to step S213. Then, the registered voice data acquisition unit 213 acquires the registered voice data for which the similarity has not been calculated from among the plurality of registered voice data stored in the registered voice data storage unit 2124.

[0329] On the other hand, if it is determined that the similarity between the voice data to be identified and all the registered voice data has been calculated (YES in step S220), in step S221, the similarity determination unit 232 determines whether the highest similarity is greater than the threshold value.

[0330] Note that the processes of steps S221 to S224 are the same as the processes of steps S41 to S44 shown in FIG. 7, and thus the description thereof is omitted.

[0331] As described above, when the utterance time of at least one of the speaker of the voice data to be identified and the speaker of the registered voice data is equal to or longer than a predetermined time, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the third speaker identification model that has been machine-learned using voice data whose utterance time is equal to or longer than the predetermined time, whereby the speaker of the voice data to be identified is identified. Also, when the utterance time of at least one of the speaker of the voice data to be identified and the speaker of the registered voice data is shorter than the predetermined time, the feature amount of the voice data to be identified and the feature amount of the registered voice data are input into the fourth speaker identification model that has been machine-learned using voice data whose utterance time is shorter than the predetermined time, whereby the speaker of the voice data to be identified is identified.

[0332] Therefore, the speaker of the voice data to be identified is identified by the third speaker identification model and the fourth speaker identification model according to the length of the speaking time of at least one of the voice data to be identified and the registered voice data, so that the accuracy of identifying whether the speaker to be identified is a pre-registered speaker can be improved.

[0333] In addition, in the fourth embodiment, the model selection unit 2174 selects one of the third speaker identification model and the fourth speaker identification model based on the speaking time associated with the registered voice data. However, the present disclosure is not particularly limited thereto. The speaker identification device 24 may include a speaking time measurement unit that measures the speaking time of the speaker of the registered voice data acquired by the registered voice data acquisition unit 213. The speaking time measurement unit may output the measured speaking time to the model selection unit 2174. When the speaking time of the speaker of the registered voice data acquired by the registered voice data acquisition unit 213 is measured, the speaking time measurement unit 207 becomes unnecessary, and the registration unit 2064 may store only the registered target voice data acquired by the registered target voice data acquisition unit 201 as the registered voice data in the registered voice data storage unit 2124.

[0334] Also, in the fourth embodiment, the model selection unit 2174 selects either the third speaker identification model or the fourth speaker identification model based on the speaking time of the speaker of the registered voice data. However, the present disclosure is not particularly limited thereto. The model selection unit 2174 may select either the third speaker identification model or the fourth speaker identification model based on the speaking time of the speaker of the voice data to be identified. In this case, the speaker identification device 24 may include a speaking time measurement unit that measures the speaking time of the speaker of the voice data to be identified. The speaking time measurement unit may measure the speaking time of the voice data to be identified acquired by the voice data acquisition unit 211 for identification and output the measured speaking time to the model selection unit 2174. When the speaking time of the speaker of the voice data to be identified is equal to or longer than a predetermined time, the model selection unit 2174 may select the third speaker identification model, and when the speaking time of the speaker of the voice data to be identified is shorter than the predetermined time, the model selection unit 2174 may select the fourth speaker identification model. Note that the predetermined time is, for example, 30 seconds. Also, when the speaking time of the speaker of the voice data to be identified is measured, the speaking time measurement unit 207 becomes unnecessary, and the registration unit 2064 may store only the voice data to be registered acquired by the voice data acquisition unit 201 for registration as registered voice data in the registered voice data storage unit 2124.

[0335] Also, in Embodiment 4, the model selection unit 2174 may select either the third speaker identification model or the fourth speaker identification model based on both the speaking time of the speaker of the registered voice data and the speaking time of the speaker of the voice data to be identified. When both the speaking time of the speaker of the registered voice data and the speaking time of the speaker of the voice data to be identified are equal to or longer than a predetermined time, the model selection unit 2174 may select the third speaker identification model. Also, when at least one of the speaking time of the speaker of the registered voice data and the speaking time of the speaker of the voice data to be identified is shorter than the predetermined time, the model selection unit 2174 may select the fourth speaker identification model. Note that the predetermined time is, for example, 20 seconds. In this case, the speaker identification device 24 may further include a speaking time measurement unit that measures the speaking time of the speaker of the voice data to be identified. Also, the speaker identification device 24 may include a speaking time measurement unit that measures the speaking time of the speaker of the registered voice data acquired by the registered voice data acquisition unit 213 without including the speaking time measurement unit 207.

[0336] Furthermore, in Embodiment 4, when either one of the speaking time of the speaker of the registered voice data and the speaking time of the speaker of the voice data to be identified is equal to or longer than a predetermined time, the model selection unit 2174 may select the third speaker identification model. Also, when both the speaking time of the speaker of the registered voice data and the speaking time of the speaker of the voice data to be identified are shorter than the predetermined time, the model selection unit 2174 may select the fourth speaker identification model. Note that the predetermined time is, for example, 100 seconds. In this case, the speaker identification device 24 may further include a speaking time measurement unit that measures the speaking time of the speaker of the voice data to be identified. Also, the speaker identification device 24 may include a speaking time measurement unit that measures the speaking time of the speaker of the registered voice data acquired by the registered voice data acquisition unit 213 without including the speaking time measurement unit 207.

[0337] Subsequently, the operation of the speaker identification model generation process of the speaker identification model generation device 44 in Embodiment 4 will be described.

[0338] FIG. 30 is a first flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device according to the fourth embodiment, and FIG. 31 is a second flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device according to the fourth embodiment.

[0339] First, in step S231, the long-duration voice data acquisition unit 422 acquires a plurality of long-duration voice data with a voice duration of a predetermined time or longer and provided with speaker identification labels for identifying a speaker from the long-duration voice data storage unit 421.

[0340] Next, in step S232, the feature extraction unit 423 extracts the features of the plurality of long-duration voice data acquired by the long-duration voice data acquisition unit 422.

[0341] Next, in step S233, the third speaker identification model generation unit 424 acquires, as teacher data, the features of the first long-duration voice data and the second long-duration voice data among the plurality of long-duration voice data, and the similarity of the speaker identification labels of the first long-duration voice data and the second long-duration voice data.

[0342] Next, in step S234, the third speaker identification model generation unit 424 uses the acquired teacher data to perform machine learning on a third speaker identification model in which the input is the features of two voice data and the output is the similarity of the two voice data.

[0343] Next, in step S235, the third speaker identification model generation unit 424 determines whether or not the third speaker identification model has been machine-learned using all combinations of the plurality of long-duration voice data. Here, if it is determined that the third speaker identification model has not been machine-learned using all combinations of the long-duration voice data (NO in step S235), the process returns to step S233. Then, the third speaker identification model generation unit 424 acquires, as teacher data, the feature amounts of the first long-duration voice data and the second long-duration voice data of the combinations not used for machine learning among the plurality of long-duration voice data, and the similarity of the speaker identification labels of the first long-duration voice data and the second long-duration voice data.

[0344] On the other hand, if it is determined that the third speaker identification model has been machine-learned using all combinations of the long-duration voice data (YES in step S235), in step S236, the third speaker identification model generation unit 424 stores the third speaker identification model generated by machine learning in the third speaker identification model storage unit 425.

[0345] Next, in step S237, the short-duration voice data acquisition unit 432 acquires, from the short-duration voice data storage unit 431, a plurality of short-duration voice data with a speaking time shorter than a predetermined time and to which speaker identification labels for identifying speakers are assigned.

[0346] Next, in step S238, the feature amount extraction unit 433 extracts the feature amounts of the plurality of short-duration voice data acquired by the short-duration voice data acquisition unit 432.

[0347] Next, in step S239, the fourth speaker identification model generation unit 434 acquires, as teacher data, the feature amounts of the first short-duration voice data and the second short-duration voice data among the plurality of short-duration voice data, and the similarity of the speaker identification labels of the first short-duration voice data and the second short-duration voice data.

[0348] Next, in step S240, the fourth speaker identification model generation unit 434 uses the acquired teacher data to perform machine learning on a fourth speaker identification model, where the input is the feature amounts of two pieces of voice data and the output is the similarity between the two pieces of voice data.

[0349] Next, in step S241, the fourth speaker identification model generation unit 434 determines whether or not it has performed machine learning on the fourth speaker identification model using combinations of all of the plurality of short-time voice data. Here, if it is determined that machine learning has not been performed on the fourth speaker identification model using combinations of all of the short-time voice data (NO in step S241), the process returns to step S239. Then, the fourth speaker identification model generation unit 434 acquires, as teacher data, the feature amounts of the first short-time voice data and the second short-time voice data of the combination not used for machine learning among the plurality of short-time voice data, and the similarity between the speaker identification labels of the first short-time voice data and the second short-time voice data.

[0350] On the other hand, if it is determined that machine learning has been performed on the fourth speaker identification model using combinations of all of the short-time voice data (YES in step S241), in step S242, the fourth speaker identification model generation unit 434 stores the fourth speaker identification model generated by machine learning in the fourth speaker identification model storage unit 435.

[0351] (Embodiment 5) In the above-described embodiment 4, the similarity between each of all of the registered voice data stored in the registered voice data storage unit 2124 and the voice data to be identified is calculated, and the speaker of the registered voice data having the highest similarity is identified as the speaker of the voice data to be identified. On the other hand, in embodiment 5, identification information of the speaker of the voice data to be identified is input, and one piece of registered voice data that has been previously associated with the identification information is acquired from among the plurality of registered voice data stored in the registered voice data storage unit 212. Then, the similarity between the one piece of registered voice data and the voice data to be identified is calculated, and if the similarity is higher than a threshold value, the speaker of the registered voice data is identified as the speaker of the voice data to be identified.

[0352] FIG. 32 is a diagram showing the configuration of the speaker identification system according to Embodiment 5 of the present disclosure.

[0353] The speaker identification system shown in FIG. 32 includes a microphone 1 and a speaker identification device 25. Note that the speaker identification device 25 may or may not include the microphone 1.

[0354] In Embodiment 5, the same components as those in Embodiments 1 to 4 are denoted by the same reference numerals, and the description thereof is omitted.

[0355] The speaker identification device 25 includes a registration target voice data acquisition unit 201, an utterance time measurement unit 207, a registration unit 2065, an identification target voice data acquisition unit 211, a registered voice data storage unit 2125, a registered voice data acquisition unit 2135, a feature amount extraction unit 214, a feature amount extraction unit 215, a speaker identification model storage unit 2164, a model selection unit 2174, a speaker identification unit 2185, an identification result output unit 219, an input reception unit 221, and an identification information acquisition unit 222.

[0356] The registration unit 2065 registers the registration target voice data in which the utterance time information indicating the utterance time measured by the utterance time measurement unit 207 and the identification information acquired by the identification information acquisition unit 222 are associated as registered voice data. The registration unit 2065 registers the registered voice data in the registered voice data storage unit 2125.

[0357] The registered voice data storage unit 2125 stores the registered voice data in which the utterance time information and the identification information are associated. The registered voice data storage unit 2125 stores a plurality of registered voice data. The plurality of registered voice data are associated with the identification information for identifying the speakers of the plurality of registered voice data respectively.

[0358] The registered voice data acquisition unit 2135 acquires the registered voice data in which the identification information matching the identification information acquired by the identification information acquisition unit 222 is associated from among the plurality of registered voice data registered in the registered voice data storage unit 2125.

[0359] The speaker identification unit 2185 includes a similarity calculation unit 2315 and a similarity determination unit 2325.

[0360] The similarity calculation unit 2315 inputs the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected third speaker identification model or the fourth speaker identification model, and obtains the similarity between the voice data to be identified and the registered voice data from either the third speaker identification model or the fourth speaker identification model.

[0361] When the obtained similarity is higher than the threshold value, the similarity determination unit 2325 identifies the speaker of the registered voice data as the speaker of the voice data to be identified.

[0362] Subsequently, the operations of the registration process and the speaker identification process of the speaker identification device 25 in the fifth embodiment will be described.

[0363] FIG. 33 is a flowchart for explaining the operation of the registration process of the speaker identification device in the fifth embodiment.

[0364] The processes of step S251 and step S252 are the same as the processes of step S121 and step S122 shown in FIG. 15, so the description thereof will be omitted. Also, the process of step S253 is the same as the process of step S202 shown in FIG. 27, so the description thereof will be omitted.

[0365] Next, in step S254, the registration unit 2065 stores the registration target voice data in which the utterance time information measured by the utterance time measurement unit 207 and the identification information obtained by the identification information acquisition unit 222 are associated as the registered voice data in the registered voice data storage unit 2125. As a result, the registered voice data storage unit 2125 stores the registered voice data in which the utterance time information and the identification information are associated.

[0366] FIG. 34 is a first flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 5, and FIG. 35 is a second flowchart for explaining the operation of the speaker identification process of the speaker identification device according to Embodiment 5.

[0367] The processes of steps S261 to S264 are the same as the processes of steps S131 to S134 shown in FIG. 16, and thus the description thereof is omitted.

[0368] Next, in step S265, the registered voice data acquisition unit 2135 acquires the registered voice data associated with the identification information that matches the identification information acquired by the identification information acquisition unit 222 from among the plurality of registered voice data registered in the registered voice data storage unit 2125.

[0369] The processes of steps S266 to S270 are the same as the processes of steps S214 to S218 shown in FIG. 28, and thus the description thereof is omitted.

[0370] Next, in step S271, the similarity calculation unit 2315 inputs the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected third speaker identification model or the fourth speaker identification model, and thereby acquires the similarity between the voice data to be identified and the registered voice data from either the third speaker identification model or the fourth speaker identification model.

[0371] Next, in step S272, the similarity determination unit 2325 determines whether the similarity calculated by the similarity calculation unit 2315 is greater than the threshold value.

[0372] Here, when it is determined that the similarity calculated by the similarity calculation unit 2315 is greater than the threshold value (YES in step S272), in step S273, the similarity determination unit 2325 identifies the speaker of the registered voice data as the speaker of the voice data to be identified.

[0373] On the other hand, when it is determined that the similarity calculated by the similarity calculation unit 2315 is equal to or less than the threshold value (NO in step S272), in step S274, the similarity determination unit 2325 determines that the speaker of the voice data to be identified is not the speaker of the registered voice data.

[0374] Since the process of step S275 is the same as the process of step S44 shown in FIG. 7, the description thereof is omitted.

[0375] As described above, in the fifth embodiment, only the similarity between the registered voice data associated with the identification information and the voice data to be identified is calculated. Therefore, compared with the fourth embodiment in which a plurality of similarities between each of the plurality of registered voice data and the voice data to be identified are calculated, in the fifth embodiment, the processing load of similarity calculation can be reduced.

[0376] (Embodiment 6) In the fourth and fifth embodiments described above, either the third speaker identification model or the fourth speaker identification model is selected according to the speaking time of the speaker of the registered voice data. However, when two different third speaker identification models and fourth speaker identification models are used, the output value ranges of the third speaker identification model and the fourth speaker identification model may be different. Therefore, in the sixth embodiment, at the time of registering voice data, a third threshold value and a fourth threshold value capable of identifying the same speaker are calculated for each of the third speaker identification model and the fourth speaker identification model, and the calculated third threshold value and fourth threshold value are stored. Further, at the time of identifying voice data, the similarity is corrected by subtracting the third threshold value or the fourth threshold value from the calculated similarity between the voice data to be identified and the registered voice data. Then, the corrected similarity is compared with a fifth threshold value common to the third speaker identification model and the fourth speaker identification model, whereby the speaker of the voice data to be identified is identified.

[0377] First, the speaker identification model generation device in the sixth embodiment of the present disclosure will be described.

[0378] FIG. 36 is a diagram showing the configuration of the speaker identification model generation device according to Embodiment 6 of the present disclosure.

[0379] The speaker identification model generation device 46 shown in FIG. 36 includes a long-duration voice data storage unit 421, a long-duration voice data acquisition unit 422, a feature quantity extraction unit 423, a third speaker identification model generation unit 424, a third speaker identification model storage unit 425, a third speaker identification unit 426, a third threshold calculation unit 427, a threshold storage unit 428, a short-duration voice data storage unit 431, a short-duration voice data acquisition unit 432, a feature quantity extraction unit 433, a fourth speaker identification model generation unit 434, a fourth speaker identification model storage unit 435, a fourth speaker identification unit 436, and a fourth threshold calculation unit 437.

[0380] In addition, in Embodiment 6, the same components as those in Embodiments 1 to 5 are denoted by the same reference numerals, and the description thereof is omitted.

[0381] The third speaker identification unit 426 inputs all combinations of the feature quantities of two pieces of voice data out of a plurality of voice data whose utterance time is equal to or longer than a predetermined time into the third speaker identification model, and thereby obtains the similarity of each of the plurality of combinations of the two pieces of voice data from the third speaker identification model.

[0382] The third threshold calculation unit 427 calculates a third threshold that can distinguish the similarity between two pieces of voice data of the same speaker and the similarity between two pieces of voice data of different speakers. The third threshold calculation unit 427 calculates the third threshold by performing regression analysis on the plurality of similarities calculated by the third speaker identification unit 426.

[0383] The fourth speaker identification unit 436 inputs all combinations of the feature quantities of two pieces of voice data out of a plurality of voice data whose utterance time is shorter than a predetermined time into the fourth speaker identification model, and thereby obtains the similarity of each of the plurality of combinations of the two pieces of voice data from the fourth speaker identification model.

[0384] The fourth threshold calculation unit 437 calculates a fourth threshold that can distinguish the similarity between two voice data of the same speaker and the similarity between two voice data of different speakers. The fourth threshold calculation unit 437 calculates the fourth threshold by performing regression analysis on a plurality of similarities calculated by the fourth speaker identification unit 436.

[0385] The threshold storage unit 428 stores the third threshold calculated by the third threshold calculation unit 427 and the fourth threshold calculated by the fourth threshold calculation unit 437.

[0386] Subsequently, the speaker identification system in Embodiment 6 of the present disclosure will be described.

[0387] FIG. 37 is a diagram showing the configuration of the speaker identification system in Embodiment 6 of the present disclosure.

[0388] The speaker identification system shown in FIG. 37 includes a microphone 1 and a speaker identification device 26. Note that the speaker identification device 26 may or may not include the microphone 1.

[0389] In addition, in this Embodiment 6, the same components as those in Embodiments 1 to 5 are denoted by the same reference numerals, and the description thereof is omitted.

[0390] The speaker identification device 26 includes a registration target voice data acquisition unit 201, a speech time measurement unit 207, a registration unit 2065, an identification target voice data acquisition unit 211, a registered voice data storage unit 2125, a registered voice data acquisition unit 2135, a feature amount extraction unit 214, a feature amount extraction unit 215, a speaker identification model storage unit 2164, a model selection unit 2174, a speaker identification unit 2186, an identification result output unit 219, an input reception unit 221, an identification information acquisition unit 222, and a threshold storage unit 2236.

[0391] The speaker identification unit 2186 includes a similarity calculation unit 2315, a similarity correction unit 2336, and a similarity determination unit 2326.

[0392] When the similarity correction unit 2336 obtains the similarity from the third speaker identification model, it subtracts the third threshold value from the obtained similarity. Also, when the similarity correction unit 2336 obtains the similarity from the fourth speaker identification model, it subtracts the fourth threshold value from the obtained similarity. When the similarity is calculated by the similarity calculation unit 2315 using the third speaker identification model, the similarity correction unit 2336 reads the third threshold value from the threshold storage unit 2236 and subtracts the third threshold value from the calculated similarity. Also, when the similarity is calculated by the similarity calculation unit 2315 using the fourth speaker identification model, the similarity correction unit 2336 reads the fourth threshold value from the threshold storage unit 2236 and subtracts the fourth threshold value from the calculated similarity.

[0393] The threshold storage unit 2236 stores in advance a third threshold value for correcting the similarity calculated using the third speaker identification model and a fourth threshold value for correcting the similarity calculated using the fourth speaker identification model. The threshold storage unit 2236 stores in advance the third threshold value and the fourth threshold value generated by the speaker identification model generation device 46.

[0394] Note that the speaker identification model generation device 46 may transmit the third threshold value and the fourth threshold value stored in the threshold storage unit 428 to the speaker identification device 26. The speaker identification device 26 may store the received third threshold value and fourth threshold value in the threshold storage unit 2236. Also, at the time of manufacturing the speaker identification device 26, the third threshold value and the fourth threshold value generated by the speaker identification model generation device 46 may be stored in the threshold storage unit 2236.

[0395] When the similarity obtained by subtracting the third threshold value or the fourth threshold value by the similarity correction unit 2336 is higher than the fifth threshold value, the similarity determination unit 2326 identifies the speaker of the registered voice data as the speaker of the identification target voice data.

[0396] Subsequently, the operation of the speaker identification model generation process of the speaker identification model generation device 46 in the sixth embodiment will be described.

[0397] FIG. 38 is a first flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 6, FIG. 39 is a second flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 6, and FIG. 40 is a third flowchart for explaining the operation of the speaker identification model generation process of the speaker identification model generation device in Embodiment 6.

[0398] The processes of steps S281 to S286 are the same as the processes of steps S231 to S236 shown in FIG. 30, so the description thereof is omitted.

[0399] Next, in step S287, the third speaker identification unit 426 acquires the third speaker identification model from the third speaker identification model storage unit 425.

[0400] Next, in step S288, the third speaker identification unit 426 acquires the feature amounts of two long-duration voice data from among the feature amounts of the plurality of long-duration voice data extracted by the feature amount extraction unit 423.

[0401] Next, in step S289, the third speaker identification unit 426 calculates the similarity between the two long-duration voice data by inputting the acquired feature amounts of the two long-duration voice data into the third speaker identification model. Note that the two long-duration voice data are either two long-duration voice data uttered by one speaker or two long-duration voice data uttered by two speakers. At this time, the similarity when the two long-duration voice data are long-duration voice data uttered by one speaker is higher than the similarity when the two long-duration voice data are long-duration voice data uttered by two speakers.

[0402] Next, in step S290, the third speaker identification unit 426 determines whether the similarity of all combinations of long-duration voice data has been calculated. Here, if it is determined that the similarity of all combinations of long-duration voice data has not been calculated (NO in step S290), the process returns to step S288. Then, the third speaker identification unit 426 acquires, from the feature amounts of a plurality of long-duration voice data, the feature amounts of two long-duration voice data for which the similarity has not been calculated, from the feature amount extraction unit 423.

[0403] On the other hand, if it is determined that the similarity of all combinations of long-duration voice data has been calculated (YES in step S290), in step S291, the third threshold calculation unit 427 calculates a third threshold that can distinguish the similarity between two long-duration voice data of the same speaker and the similarity between two long-duration voice data of different speakers, by performing regression analysis on the plurality of similarities calculated by the third speaker identification unit 426.

[0404] Next, in step S292, the third threshold calculation unit 427 stores the calculated third threshold in the threshold storage unit 428.

[0405] The processes of steps S293 to S298 are the same as the processes of steps S237 to S242 shown in FIGS. 30 and 31, and thus the description thereof is omitted.

[0406] Next, in step S299, the fourth speaker identification unit 436 acquires the fourth speaker identification model from the fourth speaker identification model storage unit 435.

[0407] Next, in step S300, the fourth speaker identification unit 436 acquires the feature amounts of two short-duration voice data from among the feature amounts of the plurality of short-duration voice data extracted by the feature amount extraction unit 433.

[0408] Next, in step S301, the fourth speaker identification unit 436 calculates the similarity of the two short-time voice data by inputting the feature amounts of the acquired two short-time voice data into the fourth speaker identification model. Note that the two short-time voice data are either two short-time voice data spoken by one speaker or two short-time voice data spoken by two speakers. At this time, the similarity when the two short-time voice data are short-time voice data spoken by one speaker is higher than the similarity when the two short-time voice data are short-time voice data spoken by two speakers.

[0409] Next, in step S302, the fourth speaker identification unit 436 determines whether the similarities of all combinations of the short-time voice data have been calculated. Here, when it is determined that the similarities of all combinations of the short-time voice data have not been calculated (NO in step S302), the process returns to step S300. Then, the fourth speaker identification unit 436 acquires, from the feature extraction unit 433, the feature amounts of two short-time voice data for which the similarity has not been calculated, from among the feature amounts of the plurality of short-time voice data.

[0410] On the other hand, when it is determined that the similarities of all combinations of the short-time voice data have been calculated (YES in step S302), in step S303, the fourth threshold calculation unit 437 calculates a fourth threshold that can distinguish the similarity between two short-time voice data of the same speaker and the similarity between two short-time voice data of different speakers, by performing a regression analysis on the plurality of similarities calculated by the fourth speaker identification unit 436.

[0411] Next, in step S304, the fourth threshold calculation unit 437 stores the calculated fourth threshold in the threshold storage unit 428.

[0412] Subsequently, the operation of the speaker identification process of the speaker identification device 26 in the sixth embodiment will be described.

[0413] FIG. 41 is a first flowchart for explaining the operation of the speaker identification process of the speaker identification device in Embodiment 6, and FIG. 42 is a second flowchart for explaining the operation of the speaker identification process of the speaker identification device in Embodiment 6.

[0414] The processes of steps S311 to S321 are the same as the processes of steps S261 to S271 shown in FIGS. 34 and 35, and thus the description thereof is omitted.

[0415] Next, in step S322, the similarity correction unit 2336 acquires the third threshold value or the fourth threshold value from the threshold value storage unit 2236. At this time, when the third speaker identification model is selected by the model selection unit 2174, the similarity correction unit 2336 acquires the third threshold value from the threshold value storage unit 2236. Also, when the fourth speaker identification model is selected by the model selection unit 2174, the similarity correction unit 2336 acquires the fourth threshold value from the threshold value storage unit 2236.

[0416] Next, in step S323, the similarity correction unit 2336 corrects the similarity calculated by the similarity calculation unit 2315 using the acquired third threshold value or fourth threshold value. At this time, the similarity correction unit 2336 subtracts the third threshold value or the fourth threshold value from the similarity calculated by the similarity calculation unit 2315.

[0417] Next, in step S324, the similarity determination unit 2326 determines whether the similarity corrected by the similarity correction unit 2336 is greater than the fifth threshold value. Note that the fifth threshold value is, for example, 0. When the corrected similarity is greater than 0, the similarity determination unit 2326 determines that the identification target voice data matches the registered voice data registered in advance, and when the corrected similarity is 0 or less, the similarity determination unit 2326 determines that the identification target voice data does not match the registered voice data registered in advance.

[0418] Here, when it is determined that the similarity corrected by the similarity correction unit 2336 is greater than the fifth threshold value (YES in step S324), in step S325, the similarity determination unit 2326 identifies the speaker of the registered voice data as the speaker of the voice data to be identified.

[0419] On the other hand, when it is determined that the similarity corrected by the similarity correction unit 2336 is less than or equal to the fifth threshold value (NO in step S324), in step S326, the similarity determination unit 2326 determines that the speaker of the voice data to be identified is not the speaker of the registered voice data.

[0420] The process of step S327 is the same as the process of step S275 shown in FIG. 35, and thus the description thereof is omitted.

[0421] When two different third speaker identification models and fourth speaker identification models are used, the output value ranges of the third speaker identification model and the fourth speaker identification model may be different. Therefore, in the sixth embodiment, at the time of registration, third threshold values and fourth threshold values capable of identifying the same speaker are calculated for each of the third speaker identification model and the fourth speaker identification model. Further, at the time of speaker identification, the similarity is corrected by subtracting the third threshold value or the fourth threshold value from the calculated similarity between the voice data to be identified and the registered voice data. Then, by comparing the corrected similarity with the fifth threshold value common to the third speaker identification model and the fourth speaker identification model, the speaker of the voice data to be identified can be identified with higher accuracy.

[0422] In each of the above embodiments, each component may be configured by dedicated hardware or may be realized by executing a software program suitable for each component. Each component may be realized by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or a semiconductor memory.

[0423] Part or all of the functions of the device according to the embodiments of the present disclosure are typically realized as LSI (Large Scale Integration), which is an integrated circuit. These may be individually formed into one chip, or may be formed into one chip so as to include part or all of them. Further, the integration is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connection and setting of circuit cells inside the LSI may be used.

[0424] Also, part or all of the functions of the device according to the embodiments of the present disclosure may be realized by a processor such as a CPU executing a program.

[0425] Also, all the numbers used above are for exemplification to specifically explain the present disclosure, and the present disclosure is not limited to the exemplified numbers.

[0426] Also, the order in which each step shown in the above flowchart is executed is for exemplification to specifically explain the present disclosure, and may be in an order other than the above as long as the same effect can be obtained. Also, a part of the above steps may be executed simultaneously (in parallel) with other steps.

Industrial Applicability

[0427] The technology according to the present disclosure can improve the accuracy of identifying whether the speaker to be identified is a pre-registered speaker, and thus is useful as a technology for identifying speakers.

Claims

1. A computer, obtains the voice data to be identified, obtains the registered voice data registered in advance, extracts the feature amount of the voice data to be identified, extracts the feature amount of the registered voice data, when either the speaker of the voice data to be identified or the speaker of the registered voice data is male, a first speaker identification model learned by machine learning using male voice data is selected to identify the male speaker, and when either the speaker of the voice data to be identified or the speaker of the registered voice data is female, a second speaker identification model learned by machine learning using female voice data is selected to identify the female speaker, by inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, the speaker of the voice data to be identified is identified, obtains the voice data to be registered, extracts the feature amount of the voice data to be registered, identifies the gender of the speaker of the voice data to be registered using the feature amount of the voice data to be registered, registers the voice data to be registered associated with the identified gender as the registered voice data, In the identification of the gender, obtains a gender identification model learned by machine learning using male and female voice data to identify the gender of the speaker, by inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data stored in advance into the gender identification model, the similarity between the voice data to be registered and each of the plurality of male voice data is obtained from the gender identification model, calculates the average of the obtained plurality of similarities as the average male similarity, By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored female voice data into the gender discrimination model, the similarity between the voice data to be registered and each of the plurality of female voice data is obtained from the gender discrimination model. Calculate the average of the obtained plurality of similarities as the average female similarity. When the average male similarity is higher than the average female similarity, identify the gender of the speaker of the voice data to be registered as male. When the average male similarity is lower than the average female similarity, identify the gender of the speaker of the voice data to be registered as female. Speaker identification method.

2. In the selection of the first speaker discrimination model or the second speaker discrimination model, When the gender of the speaker of the registered voice data is male, select the first speaker discrimination model. When the gender of the speaker of the registered voice data is female, select the second speaker discrimination model. The speaker identification method according to claim 1.

3. A computer Obtain voice data to be identified. Obtain pre-registered registered voice data. Extract the feature amount of the voice data to be identified. Extract the feature amount of the registered voice data. When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, select a first speaker discrimination model learned by machine learning using male voice data to identify the male speaker. When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, select a second speaker discrimination model learned by machine learning using female voice data to identify the female speaker. Inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model to identify the speaker of the voice data to be identified. Obtain voice data to be registered. Extract the feature amount of the voice data to be registered. Using the feature amount of the voice data to be registered, identify the gender of the speaker of the voice data to be registered. Register the voice data to be registered associated with the identified gender as the registered voice data. In the identification of the gender, Obtain a gender identification model that is machine-learned using voice data of men and women to identify the gender of the speaker. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored voice data of men into the gender identification model, obtain the similarity between the voice data to be registered and each of the plurality of voice data of men from the gender identification model. Calculate the maximum value among the obtained plurality of similarities as the maximum male similarity. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of pre-stored voice data of women into the gender identification model, obtain the similarity between the voice data to be registered and each of the plurality of voice data of women from the gender identification model. Calculate the maximum value among the obtained plurality of similarities as the maximum female similarity. If the maximum male similarity is higher than the maximum female similarity, identify the gender of the speaker of the voice data to be registered as male. If the maximum male similarity is lower than the maximum female similarity, identify the gender of the speaker of the voice data to be registered as female. Speaker identification method.

4. The computer Obtain voice data to be identified. Obtain pre-registered registered voice data. Extract the feature quantities of the voice data to be identified, Extract the feature quantities of the registered voice data, When either the speaker of the voice data to be identified or the speaker of the registered voice data is male, select the first speaker identification model that has been machine-learned using male voice data to identify the male speaker. When either the speaker of the voice data to be identified or the speaker of the registered voice data is female, select the second speaker identification model that has been machine-learned using female voice data to identify the female speaker. Input the feature quantities of the voice data to be identified and the feature quantities of the registered voice data into either the selected first speaker identification model or the second speaker identification model to identify the speaker of the voice data to be identified. The registered voice data includes a plurality of registered voice data, In the identification of the speaker, Input the feature quantities of the voice data to be identified and the feature quantities of each of the plurality of registered voice data into either the selected first speaker identification model or the second speaker identification model to obtain the similarity between the voice data to be identified and each of the plurality of registered voice data from either the first speaker identification model or the second speaker identification model. Identify the speaker of the registered voice data with the highest obtained similarity as the speaker of the voice data to be identified. Speaker identification method.

5. The registered voice data includes a plurality of registered voice data, The plurality of registered voice data is associated with identification information for identifying the speaker of each of the plurality of registered voice data. Furthermore, obtain identification information for identifying the speaker of the voice data to be identified. In the acquisition of the registered voice data, acquire the registered voice data with which the identification information that matches the acquired identification information is associated from among the plurality of registered voice data. In the identification of the speaker, By inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, the similarity between the voice data to be identified and the registered voice data is obtained from either the first speaker identification model or the second speaker identification model, When the obtained similarity is higher than the threshold value, the speaker of the registered voice data is identified as the speaker of the voice data to be identified. The speaker identification method according to any one of claims 1 to 3.

6. During machine learning, all combinations of the feature amounts of two pieces of voice data among a plurality of male voice data are input into the first speaker identification model, whereby the similarity of each of the plurality of combinations of the two pieces of voice data is obtained from the first speaker identification model, and a first threshold value capable of distinguishing the similarity of the two pieces of voice data of the same speaker from the similarity of the two pieces of voice data of different speakers is calculated. During machine learning, all combinations of the feature amounts of two pieces of voice data among a plurality of female voice data are input into the second speaker identification model, whereby the similarity of each of the plurality of combinations of the two pieces of voice data is obtained from the second speaker identification model, and a second threshold value capable of distinguishing the similarity of the two pieces of voice data of the same speaker from the similarity of the two pieces of voice data of different speakers is calculated. In the identification of the speaker, when the similarity is obtained from the first speaker identification model, the first threshold value is subtracted from the obtained similarity, and when the similarity is obtained from the second speaker identification model, the second threshold value is subtracted from the obtained similarity. The speaker identification method according to claim 4.

7. A computer obtains voice data to be identified, obtains registered voice data registered in advance, extracts the feature amount of the voice data to be identified, extracts the feature amount of the registered voice data, When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, a first speaker identification model that has been machine-learned using male voice data is selected to identify the male speaker. When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, a second speaker identification model that has been machine-learned using female voice data is selected to identify the female speaker. By inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, the speaker of the voice data to be identified is identified. The registered voice data includes a plurality of registered voice data. The plurality of registered voice data is associated with identification information for identifying the speaker of each of the plurality of registered voice data. Furthermore, identification information for identifying the speaker of the voice data to be identified is obtained. In the acquisition of the registered voice data, registered voice data associated with identification information that matches the obtained identification information is acquired from among the plurality of registered voice data. In the identification of the speaker By inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, the similarity between the voice data to be identified and the registered voice data is obtained from either the first speaker identification model or the second speaker identification model. When the obtained similarity is higher than a threshold value, the speaker of the registered voice data is identified as the speaker of the voice data to be identified. During machine learning, all combinations of the feature amounts of two voice data among a plurality of male voice data are input into the first speaker identification model, whereby the similarity of each of the plurality of combinations of the two voice data is obtained from the first speaker identification model, and a first threshold value capable of distinguishing the similarity of the two voice data of the same speaker from the similarity of the two voice data of different speakers is calculated. During machine learning, all combinations of the feature quantities of two pieces of voice data out of a plurality of voice data of women are input into the second speaker identification model, whereby the similarity of each of the plurality of combinations of the two pieces of voice data is obtained from the second speaker identification model, and a second threshold value capable of distinguishing the similarity of the two pieces of voice data of the same speaker from the similarity of the two pieces of voice data of different speakers is calculated. In the identification of the speaker, when the similarity is obtained from the first speaker identification model, the first threshold value is subtracted from the obtained similarity, and when the similarity is obtained from the second speaker identification model, the second threshold value is subtracted from the obtained similarity. Speaker identification method.

8. An identification target voice data acquisition unit that acquires identification target voice data, A registered voice data acquisition unit that acquires registered voice data registered in advance, A first extraction unit that extracts the feature quantity of the identification target voice data, A second extraction unit that extracts the feature quantity of the registered voice data, When either the gender of the speaker of the identification target voice data or the gender of the speaker of the registered voice data is male, a first speaker identification model that is machine-learned using male voice data to identify a male speaker is selected, and when either the gender of the speaker of the identification target voice data or the gender of the speaker of the registered voice data is female, a second speaker identification model that is machine-learned using female voice data to identify a female speaker is selected. A speaker identification model selection unit, By inputting the feature quantity of the identification target voice data and the feature quantity of the registered voice data into either the selected first speaker identification model or the second speaker identification model, a speaker identification unit that identifies the speaker of the identification target voice data, A registration target voice data acquisition unit that acquires registration target voice data, A third extraction unit that extracts the feature quantity of the registration target voice data, A gender identification unit that identifies the gender of the speaker of the voice data to be registered using the feature amount of the voice data to be registered; A registration unit that registers the voice data to be registered associated with the identified gender as the registered voice data; comprising; The gender identification unit: obtains a gender identification model that has been machine-learned using voice data of men and women to identify the gender of the speaker; By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of voice data of men stored in advance into the gender identification model, the similarity between the voice data to be registered and each of the plurality of voice data of men is obtained from the gender identification model; calculates the average of the obtained plurality of similarities as the average male similarity; By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of voice data of women stored in advance into the gender identification model, the similarity between the voice data to be registered and each of the plurality of voice data of women is obtained from the gender identification model; calculates the average of the obtained plurality of similarities as the average female similarity; When the average male similarity is higher than the average female similarity, the gender of the speaker of the voice data to be registered is identified as male; When the average male similarity is lower than the average female similarity, the gender of the speaker of the voice data to be registered is identified as female. Speaker identification device.

9. obtains identification target voice data; obtains registered voice data that has been registered in advance; extracts the feature amount of the identification target voice data; extracts the feature amount of the registered voice data; When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is male, a first speaker identification model that has been machine-learned using male voice data is selected to identify the male speaker. When the gender of either the speaker of the voice data to be identified or the speaker of the registered voice data is female, a second speaker identification model that has been machine-learned using female voice data is selected to identify the female speaker. By inputting the feature amount of the voice data to be identified and the feature amount of the registered voice data into either the selected first speaker identification model or the second speaker identification model, the speaker of the voice data to be identified is identified. Acquire voice data to be registered. Extract the feature amount of the voice data to be registered. Using the feature amount of the voice data to be registered, identify the gender of the speaker of the voice data to be registered. Cause the computer to function so as to register the voice data to be registered associated with the identified gender as the registered voice data. In the identification of the gender, Obtain a gender identification model that has been machine-learned using male and female voice data to identify the gender of the speaker. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of male voice data stored in advance into the gender identification model, obtain the similarity between the voice data to be registered and each of the plurality of male voice data from the gender identification model. Calculate the average of the obtained plurality of similarities as the average male similarity. By inputting the feature amount of the voice data to be registered and the feature amounts of a plurality of female voice data stored in advance into the gender identification model, obtain the similarity between the voice data to be registered and each of the plurality of female voice data from the gender identification model. Calculate the average of the obtained plurality of similarities as the average female similarity. When the average male similarity is higher than the average female similarity, identify the gender of the speaker of the voice data to be registered as male, When the average male similarity is lower than the average female similarity, identify the gender of the speaker of the voice data to be registered as female, Speaker identification program.

Citation Information

Patent Citations

  • Speaker recognition device, speaker recognition method, and speaker recognition program

    JP2014048534A

  • Voiceprint authentication processing method and device

    JP2018508799A