Speaker recognition method, recognition device and program product

By selecting the corresponding speaker recognition model according to gender, the problem of insufficient recognition accuracy in the prior art is solved, and higher recognition accuracy and processing efficiency are achieved.

CN115315746BActive Publication Date: 2025-08-22PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180023079.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2021-02-15
Publication Date
2025-08-22
Estimated Expiration
2041-02-15

AI Technical Summary

Technical Problem

In the prior art, further improvement is needed to improve the accuracy of identifying whether the speaker of the identified object is a pre-registered speaker.

Method used

By obtaining the speech data feature quantity of the identified object and the registered object, and selecting the corresponding male or female speaker recognition model according to the gender for identification, the speech data is identified using the speaker recognition model specifically for each gender.

Benefits of technology

The accuracy of identifying whether the speaker of the identified object is a pre-registered speaker is improved, and the processing load for the speaker's identification is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115315746B_ABST
    Figure CN115315746B_ABST
Patent Text Reader

Abstract

A speaker recognition device: obtains recognition target speech data; obtains registered speech data; when the gender of either the speaker of the recognition target speech data or the speaker of the registered speech data is male, selects a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers; when the gender of either the speaker of the recognition target speech data or the speaker of the registered speech data is female, selects a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers; and, by inputting feature quantities of the recognition target speech data and feature quantities of the registered speech data into either the selected first speaker recognition model or the second speaker recognition model, recognizes the speaker of the recognition target speech data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology for identifying a speaker. Background Art

[0002] Conventionally, there are known technologies for obtaining speech data of a speaker to be recognized and, based on the obtained speech data, determining whether the speaker to be recognized is a pre-registered speaker. In conventional speaker recognition, the similarity between the feature values ​​of the speech data of the target speaker and the feature values ​​of the speech data of the registered speaker is calculated. If the calculated similarity exceeds a threshold, it is determined that the target speaker and the registered speaker are the same person.

[0003] For example, Non-Patent Document 1 discloses a technique for using a speaker-specific feature quantity called an i-vector as a highly accurate feature quantity for speaker recognition.

[0004] Furthermore, for example, a technique for using x-vectors as feature quantities instead of i-vectors is disclosed in Non-Patent Document 2. X-vectors are feature quantities extracted by inputting speech data into a deep neural network generated by deep learning.

[0005] However, in the conventional technology, further improvement is required to improve the accuracy of identifying whether the speaker to be identified is a pre-registered speaker.

[0006] Prior art literature

[0007] Non-patent literature

[0008] Non-patent literature 1: Najim Dehak, Patrick Kenny, Reda Dehak, Pierre Dumouchel, Pierre Ouellet, "Front-End Factor Analysis For Speaker Verification", IEEE Transactions on Audio, Speech and Language Processing, 2011

[0009] Non-Patent Literature 2: David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povery, Sanjeev Khudanpur, "X-Vectors: Robust DNN Embeddings for Speaker Recognition", IEEE, ICASSP, 2018 Summary of the Invention

[0010] The present invention has been made to solve the above-mentioned problem, and an object of the present invention is to provide a technique capable of improving the accuracy of identifying whether a speaker to be identified is a pre-registered speaker.

[0011] One aspect of the present invention relates to a speaker recognition method in which a computer executes the following steps: obtaining speech data to be recognized; obtaining pre-registered registration speech data; extracting feature quantities of the speech data to be recognized; extracting feature quantities of the registration speech data; selecting a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers when the gender of either the speaker of the speech data to be recognized or the speaker of the registration speech data is male, and selecting a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers when the gender of either the speaker of the speech data to be recognized or the speaker of the registration speech data is female; and recognizing the speaker of the speech data to be recognized by inputting the feature quantities of the speech data to be recognized and the feature quantities of the registration speech data into either the selected first speaker recognition model or the selected second speaker recognition model.

[0012] According to the present invention, it is possible to improve the accuracy of identifying whether a speaker to be identified is a pre-registered speaker. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 1 of the present invention.

[0014] Figure 2 This is a diagram showing the configuration of a gender identification model generation device according to the first embodiment of the present invention.

[0015] Figure 3 This is a diagram showing the configuration of a speaker recognition model generation device according to Embodiment 1 of the present invention.

[0016] Figure 4 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the first embodiment of the present invention.

[0017] Figure 5 Is used to illustrate Figure 4 Flowchart of the operation of the gender recognition processing of step S3.

[0018] Figure 6 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the first embodiment.

[0019] Figure 7 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition apparatus according to the first embodiment.

[0020] Figure 8 This is a flowchart for explaining the operation of the gender identification model generation process performed by the gender identification model generation device according to the first embodiment.

[0021] Figure 9 This is a first flowchart for explaining the operation of speaker recognition model generation processing by the speaker recognition model generation device according to the first embodiment.

[0022] Figure 10 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the first embodiment.

[0023] Figure 11 1 is a diagram showing speaker recognition performance evaluation results of a conventional speaker recognition apparatus and speaker recognition performance evaluation results of the speaker recognition apparatus according to Embodiment 1. FIG.

[0024] Figure 12 This is a flowchart for explaining the operation of the gender identification process according to the first modification of the first embodiment.

[0025] Figure 13 This is a flowchart for explaining the operation of the gender identification process according to the second modification of the first embodiment.

[0026] Figure 14 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 2 of the present invention.

[0027] Figure 15 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the second embodiment.

[0028] Figure 16 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the second embodiment.

[0029] Figure 17 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the second embodiment.

[0030] Figure 18 This is a diagram showing the configuration of a speaker recognition model generation device according to a third embodiment of the present invention.

[0031] Figure 19 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 3 of the present invention.

[0032] Figure 20This is a first flowchart for explaining the operation of speaker recognition model generation processing by the speaker recognition model generation device according to the third embodiment.

[0033] Figure 21 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the third embodiment.

[0034] Figure 22 This is a third flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the third embodiment.

[0035] Figure 23 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the third embodiment.

[0036] Figure 24 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the third embodiment.

[0037] Figure 25 This is a diagram showing the configuration of a speaker recognition system according to a fourth embodiment of the present invention.

[0038] Figure 26 This is a diagram showing the configuration of a speaker recognition model generation device according to a fourth embodiment of the present invention.

[0039] Figure 27 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the fourth embodiment.

[0040] Figure 28 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the fourth embodiment.

[0041] Figure 29 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the fourth embodiment.

[0042] Figure 30 This is a first flowchart for explaining the operation of speaker recognition model generation processing by the speaker recognition model generation device according to the fourth embodiment.

[0043] Figure 31 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the fourth embodiment.

[0044] Figure 32 This is a diagram showing the configuration of a speaker recognition system according to a fifth embodiment of the present invention.

[0045] Figure 33This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the fifth embodiment.

[0046] Figure 34 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the fifth embodiment.

[0047] Figure 35 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the fifth embodiment.

[0048] Figure 36 This is a diagram showing the configuration of a speaker recognition model generation device according to a sixth embodiment of the present invention.

[0049] Figure 37 This is a diagram showing the configuration of a speaker recognition system according to a sixth embodiment of the present invention.

[0050] Figure 38 This is a first flowchart for explaining the operation of speaker recognition model generation processing by the speaker recognition model generation device according to the sixth embodiment.

[0051] Figure 39 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the sixth embodiment.

[0052] Figure 40 This is a third flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the sixth embodiment.

[0053] Figure 41 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the sixth embodiment.

[0054] Figure 42 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the sixth embodiment. DETAILED DESCRIPTION

[0055] (Basic knowledge of the present invention)

[0056] The distribution of feature quantities in speech data differs between men and women when speaking. However, conventional speaker recognition devices use a speaker recognition model generated from speech data collected regardless of the speaker's gender to identify the speaker of the target speech data. As mentioned above, conventional speaker recognition technology specifically tailored for male and female speakers has not been developed. Therefore, it is possible to improve speaker recognition accuracy by performing gender-sensitive speaker recognition.

[0057] In order to solve the above problems, one aspect of the present invention relates to a speaker recognition method that allows a computer to perform the following steps: obtaining recognition object voice data; obtaining pre-registered registration voice data; extracting feature quantities of the recognition object voice data; extracting feature quantities of the registration voice data; when the gender of either the speaker of the recognition object voice data or the speaker of the registration voice data is male, selecting a first speaker recognition model that has been machine-learned using male voice data for recognizing male speakers; when the gender of either the speaker of the recognition object voice data or the speaker of the registration voice data is female, selecting a second speaker recognition model that has been machine-learned using female voice data for recognizing female speakers; and, by inputting the feature quantities of the recognition object voice data and the feature quantities of the registration voice data into either the selected first speaker recognition model or the second speaker recognition model, recognizing the speaker of the recognition object voice data.

[0058] According to this configuration, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a second speaker recognition model generated for female identification.

[0059] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model and the second speaker recognition model specifically for each gender, thereby improving the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker.

[0060] In addition, in the speaker recognition method, when selecting the first speaker recognition model or the second speaker recognition model, when the gender of the speaker in the registered voice data is male, the first speaker recognition model may be selected, and when the gender of the speaker in the registered voice data is female, the second speaker recognition model may be selected.

[0061] With this configuration, either the first or second speaker recognition model is selected based on the gender of the speaker in the registered voice data. Therefore, the gender of the speaker in the registered voice data is identified once beforehand during registration, eliminating the need to identify the gender of the target voice data each time speaker recognition is performed, thereby reducing the processing load of speaker recognition.

[0062] In addition, in the speaker recognition method, voice data of the registration object can also be obtained; feature quantities of the voice data of the registration object can also be extracted; the gender of the speaker of the voice data of the registration object can also be identified using the feature quantities of the voice data of the registration object; and the voice data of the registration object corresponding to the identified gender can also be registered as the registration voice data.

[0063] According to this configuration, during registration, registration voice data corresponding to gender can be pre-registered, and during speaker recognition, either the first speaker recognition model or the second speaker recognition model can be easily selected using the gender corresponding to the pre-registered registration voice data.

[0064] In addition, in the speaker recognition method, when identifying the gender, a gender recognition model that has been machine-learned using male and female voice data to identify the gender of the speaker can be obtained, and the gender of the speaker of the registration object voice data can be identified by inputting the feature value of the registration object voice data into the gender recognition model.

[0065] According to this configuration, the gender of the speaker of the registration object voice data can be easily identified simply by inputting the registration object voice data into a gender identification model, which is a recognition model that performs machine learning using male and female voice data to identify the gender of the speaker.

[0066] In addition, in the speaker recognition method, when identifying the gender, the characteristic amount of the registration object voice data and the characteristic amounts of multiple pre-stored male voice data are input into the gender recognition model, the similarity between the registration object voice data and each voice data of the multiple male voice data is obtained from the gender recognition model, and the average of the obtained multiple similarities is calculated as the average male similarity. The characteristic amount of the registration object voice data and the characteristic amounts of multiple pre-stored female voice data are input into the gender recognition model, the similarity between the registration object voice data and each voice data of the multiple female voice data is obtained from the gender recognition model, and the average of the obtained multiple similarities is calculated as the average female similarity. If the average male similarity is higher than the average female similarity, the gender of the speaker of the registration object voice data is identified as male, and if the average male similarity is lower than the average female similarity, the gender of the speaker of the registration object voice data is identified as female.

[0067] According to this configuration, if the speaker of the registration target voice data is male, the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data is higher than the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data. Furthermore, if the speaker of the registration target voice data is female, the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data is higher than the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data. Therefore, by comparing the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data with the average similarity between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data, the gender of the speaker of the registration target voice data can be easily identified.

[0068] In addition, in the speaker recognition method, when identifying the gender, the characteristic amount of the registration object voice data and the characteristic amounts of multiple pre-stored male voice data are input into the gender recognition model, the similarity between the registration object voice data and each voice data of the multiple male voice data is obtained from the gender recognition model, and the maximum value of the multiple similarities obtained is calculated as the maximum male similarity. The characteristic amount of the registration object voice data and the characteristic amounts of multiple pre-stored female voice data are input into the gender recognition model, the similarity between the registration object voice data and each voice data of the multiple female voice data is obtained from the gender recognition model, and the maximum value of the multiple similarities obtained is calculated as the maximum female similarity. When the maximum male similarity is higher than the maximum female similarity, the gender of the speaker of the registration object voice data is identified as male, and when the maximum male similarity is lower than the maximum female similarity, the gender of the speaker of the registration object voice data is identified as female.

[0069] According to this configuration, if the speaker of the registration target voice data is male, the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data is higher than the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data. Furthermore, if the speaker of the registration target voice data is female, the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data is higher than the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data. Therefore, by comparing the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple male voice data with the maximum similarity among the multiple similarities between the feature quantity of the registration target voice data and the feature quantities of multiple female voice data, the gender of the speaker of the registration target voice data can be easily identified.

[0070] In addition, in the speaker recognition method, when identifying the gender, the average feature value of multiple pre-stored male voice data can be calculated, and the first similarity between the registration object voice data and the male voice data group is obtained from the gender recognition model by inputting the feature value of the registration object voice data and the average feature value of the multiple male voice data into the gender recognition model. The average feature value of multiple pre-stored female voice data is calculated, and the second similarity between the registration object voice data and the female voice data group is obtained from the gender recognition model by inputting the feature value of the registration object voice data and the average feature value of the multiple female voice data into the gender recognition model. When the first similarity is higher than the second similarity, the gender of the speaker of the registration object voice data is identified as male, and when the first similarity is lower than the second similarity, the gender of the speaker of the registration object voice data is identified as female.

[0071] According to this configuration, when the speaker of the voice data to be registered is male, the first similarity between the feature quantity of the voice data to be registered and the average feature quantity of the voice data of multiple males is higher than the second similarity between the feature quantity of the voice data to be registered and the average feature quantity of the voice data of multiple females. Furthermore, when the speaker of the voice data to be registered is female, the second similarity between the feature quantity of the voice data to be registered and the average feature quantity of the voice data of multiple females is higher than the first similarity between the feature quantity of the voice data to be registered and the average feature quantity of the voice data of multiple males. Therefore, by comparing the first similarity between the feature quantity of the voice data to the average feature quantity of the voice data of multiple males and the second similarity between the feature quantity of the voice data to the average feature quantity of the voice data of multiple females, the gender of the speaker of the voice data to be registered can be easily identified.

[0072] In addition, in the speaker recognition method, the registration voice data may also include multiple registration voice data. When identifying the speaker, the feature value of the recognition target voice data and each feature value of the multiple registration voice data are input into any one of the selected first speaker recognition model and the second speaker recognition model, and the similarity between the recognition target voice data and each registration voice data of the multiple registration voice data is obtained from any one of the first speaker recognition model and the second speaker recognition model, and the speaker of the registration voice data with the highest similarity is identified as the speaker of the recognition target voice data.

[0073] According to this configuration, the similarity between the target speech data and each of the plurality of registered speech data is obtained from either the first speaker recognition model or the second speaker recognition model, and the speaker of the registered speech data with the highest similarity is identified as the speaker of the target speech data. Therefore, the speaker of the most similar registered speech data among the plurality of registered speech data can be identified as the speaker of the target speech data.

[0074] In addition, in the speaker recognition method, the registration voice data may also include multiple registration voice data, and the multiple registration voice data correspond to the recognition information of the speaker of each registration voice data used to identify the multiple registration voice data, and the recognition information of the speaker of the recognition target voice data is also obtained; when obtaining the registration voice data, the registration voice data corresponding to the recognition information consistent with the obtained recognition information is obtained from the multiple registration voice data; when identifying the speaker, the similarity between the recognition target voice data and the registration voice data is obtained from any one of the first speaker recognition model and the second speaker recognition model selected by inputting the feature value of the recognition target voice data and the feature value of the registration voice data into any one of the selected first speaker recognition model and the second speaker recognition model; when the obtained similarity is higher than a threshold value, the speaker of the registration voice data is identified as the speaker of the recognition target voice data.

[0075] With this configuration, a single piece of registration voice data corresponding to identification information that matches the identification information used to identify the speaker of the target voice data is obtained from the plurality of registration voice data. Therefore, it is not necessary to calculate the similarity between the feature quantities of all the registration voice data and the feature quantities of the target voice data. It is sufficient to calculate the similarity between the feature quantities of only one piece of the registration voice data and the feature quantities of the target voice data. This reduces the processing load for speaker recognition.

[0076] Furthermore, in the speaker recognition method, during machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data may be input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data may be obtained from the first speaker recognition model. A first threshold value for distinguishing the similarities of the two speech data from the same speaker and the similarities of the two speech data from different speakers may be calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of female speech data may be input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data may be obtained from the second speaker recognition model. A second threshold value for distinguishing the similarities of the two speech data from the same speaker and the similarities of the two speech data from different speakers may be calculated. When recognizing the speaker, if the similarities are obtained from the first speaker recognition model, the first threshold value is subtracted from the obtained similarities, and if the similarities are obtained from the second speaker recognition model, the second threshold value is subtracted from the obtained similarities.

[0077] When using two different first and second speaker recognition models, the output value ranges of the first and second speaker recognition models may differ. Therefore, during registration, a first threshold and a second threshold that enable identification of the same speaker are calculated for each of the first and second speaker recognition models. Furthermore, when identifying a speaker, the similarity between the calculated speech data to be recognized and the registered speech data is corrected by subtracting the first or second threshold from the calculated similarity. Furthermore, by comparing this corrected similarity with a common threshold for both the first and second speaker recognition models, the speaker of the speech data to be recognized can be identified with higher accuracy.

[0078] Another aspect of the present invention relates to a speaker recognition device, comprising: a recognition target speech data acquisition unit for acquiring recognition target speech data; a registration speech data acquisition unit for acquiring pre-registered registration speech data; a first extraction unit for extracting feature quantities of the recognition target speech data; a second extraction unit for extracting feature quantities of the registration speech data; a speaker recognition model selection unit for selecting, when either the gender of the speaker of the recognition target speech data or the speaker of the registration speech data is male, a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers; and, when either the gender of the speaker of the recognition target speech data or the speaker of the registration speech data is female, a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers; and a speaker recognition unit for recognizing the speaker of the recognition target speech data by inputting the feature quantities of the recognition target speech data and the feature quantities of the registration speech data into either the selected first speaker recognition model or the selected second speaker recognition model.

[0079] According to this configuration, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a second speaker recognition model generated for female identification.

[0080] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model and the second speaker recognition model specifically for each gender, thereby improving the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker.

[0081] Yet another aspect of the present invention relates to a speaker recognition program that enables a computer to perform the following functions: obtain speech data to be recognized; obtain pre-registered registration speech data; extract feature quantities of the speech data to be recognized; extract feature quantities of the registration speech data; if the gender of either the speaker of the speech data to be recognized or the speaker of the registration speech data is male, select a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers, and if the gender of either the speaker of the speech data to be recognized or the speaker of the registration speech data is female, select a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers; and recognize the speaker of the speech data to be recognized by inputting the feature quantities of the speech data to be recognized and the feature quantities of the registration speech data into either the selected first speaker recognition model or the second speaker recognition model.

[0082] According to this configuration, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a second speaker recognition model generated for female identification.

[0083] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model and the second speaker recognition model specifically for each gender, thereby improving the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker.

[0084] Another aspect of the present invention relates to a method for generating a gender recognition model, wherein a computer executes the following steps: obtaining a plurality of speech data that are assigned gender labels indicating whether the speaker is male or female; and generating a gender recognition model through machine learning, wherein the gender recognition model uses the respective feature quantities of the first speech data and the second speech data in the plurality of speech data and the similarity of the respective gender labels of the first speech data and the second speech data as teacher data, uses the respective feature quantities of the two speech data as input, and uses the similarity of the two speech data as output.

[0085] According to this configuration, by inputting the feature values ​​of the registration voice data or the recognition target voice data and the feature values ​​of the male voice data into a gender recognition model generated by machine learning, a first similarity between the two voice data is output. Furthermore, by inputting the feature values ​​of the registration voice data or the recognition target voice data and the feature values ​​of the female voice data into the gender recognition model, a second similarity between the two voice data is output. Furthermore, by comparing the first and second similarities, the gender of the speaker of the registration voice data or the recognition target voice data can be easily estimated.

[0086] Another aspect of the present invention relates to a method for generating a speaker recognition model, wherein a computer executes the following steps: obtaining a plurality of male voice data assigned speaker recognition labels for identifying male speakers; generating a first speaker recognition model through machine learning, wherein the first speaker recognition model uses the respective feature quantities of the first male voice data and the second male voice data in the plurality of male voice data and the similarity of the respective speaker recognition labels of the first male voice data and the second male voice data as teacher data, uses the respective feature quantities of the two voice data as input, and uses the similarity of the two voice data as output; obtaining a plurality of female voice data assigned speaker recognition labels for identifying female speakers; and generating a second speaker recognition model through machine learning, wherein the second speaker recognition model uses the respective feature quantities of the first female voice data and the second female voice data in the plurality of female voice data and the similarity of the respective speaker recognition labels of the first female voice data and the second female voice data as teacher data, uses the respective feature quantities of the two voice data as input, and uses the similarity of the two voice data as output.

[0087] According to this configuration, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, if either the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is identified by inputting the feature value of the speech data to be recognized and the feature value of the registered speech data into a second speaker recognition model generated for female identification.

[0088] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model specifically for males and the second speaker recognition model specifically for females. Therefore, the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker can be improved.

[0089] Hereinafter, an embodiment of the present invention will be described with reference to the accompanying drawings. Note that the following embodiment is an example of a specific embodiment of the present invention and does not limit the technical scope of the present invention.

[0090] (Implementation Method 1)

[0091] Figure 1 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 1 of the present invention.

[0092] Figure 1 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 2. The speaker recognition device 2 may or may not include the microphone 1.

[0093] Microphone 1 picks up the sound of a speaker, converts it into voice data, and outputs it to speaker recognition device 2. When voice data is pre-registered, microphone 1 outputs the registered voice data uttered by the speaker to speaker recognition device 2. Furthermore, when identifying a speaker, microphone 1 outputs the recognized voice data uttered by the speaker to speaker recognition device 2.

[0094] The speaker recognition device 2 includes a registration object speech data acquisition unit 201, a feature value extraction unit 202, a gender recognition model storage unit 203, a gender recognition speech data storage unit 204, a gender recognition unit 205, a registration unit 206, a recognition object speech data acquisition unit 211, a registration speech data storage unit 212, a registration speech data acquisition unit 213, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 216, a model selection unit 217, a speaker recognition unit 218 and a recognition result output unit 219.

[0095] The registration target speech data acquisition unit 201, feature extraction unit 202, gender recognition unit 205, registration unit 206, recognition target speech data acquisition unit 211, registration speech data acquisition unit 213, feature extraction unit 214, feature extraction unit 215, model selection unit 217, speaker recognition unit 218, and recognition result output unit 219 are implemented by a processor. The processor is formed, for example, by a CPU (central processing unit).

[0096] Gender recognition model storage unit 203, gender recognition speech data storage unit 204, registered speech data storage unit 212, and speaker recognition model storage unit 216 are implemented using memory. The memory is formed, for example, by ROM (Read Only Memory) or EEPROM (Electrically Erasable Programmable Read Only Memory).

[0097] In addition, the speaker recognition device 2 may be, for example, a computer, a smartphone, a tablet computer, or a server.

[0098] The registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1 .

[0099] The feature extraction unit 202 extracts features from the target speech data acquired by the target speech data acquisition unit 201. An example of a feature is an i-vector. An i-vector is a low-dimensional vector feature extracted from speech data by applying factor analysis to a GMM (Gaussian Mixture Model) supervector. Since the i-vector extraction method is conventional, a detailed description is omitted. Furthermore, the feature is not limited to an i-vector; other features, such as an x-vector, may also be used.

[0100] The gender recognition model storage unit 203 stores a gender recognition model in advance. The gender recognition model is a recognition model that is machine-learned using male and female voice data to recognize the gender of a speaker. The method for generating the gender recognition model will be described later.

[0101] The gender identification speech data storage unit 204 pre-stores feature quantities of gender identification speech data used to identify the gender of the speaker in the registered speech data. The gender identification speech data includes multiple sets of male and female speech data. While the gender identification speech data storage unit 204 pre-stores feature quantities of the gender identification speech data, the present invention is not particularly limited to this, and gender identification speech data may also be pre-stored. In this case, the speaker recognition device 2 includes a feature quantity extraction unit that extracts feature quantities of the gender identification speech data.

[0102] The gender recognition unit 205 uses the features of the target speech data extracted by the feature extraction unit 202 to identify the gender of the speaker in the target speech data. The gender recognition unit 205 obtains from the gender recognition model storage unit 203 a gender recognition model that has been machine-learned using male and female speech data to identify the gender of the speaker. The gender recognition unit 205 inputs the features of the target speech data into the gender recognition model to identify the gender of the speaker in the target speech data.

[0103] Gender recognition unit 205 inputs the feature values ​​of the registration target voice data and the feature values ​​of each of the plurality of male voice data previously stored in gender recognition voice data storage unit 204 into a gender recognition model. The model then obtains the similarity between the registration target voice data and each of the plurality of male voice data. Gender recognition unit 205 then calculates the average of the obtained plurality of similarities as the average male similarity.

[0104] Furthermore, the gender recognition unit 205 inputs the feature value of the registration target voice data and each feature value of the plurality of female voice data previously stored in the gender recognition voice data storage unit 204 into the gender recognition model. The gender recognition unit 205 then obtains the similarity between the registration target voice data and each of the plurality of female voice data from the gender recognition model. The gender recognition unit 205 then calculates the average of the obtained plurality of similarities as the average female similarity.

[0105] If the average male similarity is higher than the average female similarity, the gender identification unit 205 identifies the gender of the speaker of the target voice data as male. On the other hand, if the average male similarity is lower than the average female similarity, the gender identification unit 205 identifies the gender of the speaker of the target voice data as female. Furthermore, if the average male similarity is equal to the average female similarity, the gender identification unit 205 may identify the gender of the speaker of the target voice data as male or female.

[0106] The registration unit 206 registers, as registration voice data, the voice data to be registered that corresponds to the gender information identified by the gender identification unit 205. The registration unit 206 registers the registration voice data in the registration voice data storage unit 212.

[0107] Furthermore, the speaker recognition device 2 may further include an input receiving unit for receiving information about the speaker of the voice data to be registered. Furthermore, the registration unit 206 may register the registered voice data and the speaker information in the registered voice data storage unit 212 in association with each other. The speaker information may include, for example, the speaker's name.

[0108] The recognition target speech data acquisition unit 211 acquires the recognition target speech data output from the microphone 1 .

[0109] The registered voice data storage unit 212 stores the registered voice data corresponding to the gender information. The registered voice data storage unit 212 stores a plurality of registered voice data.

[0110] The registered voice data acquisition unit 213 acquires the registered voice data registered in the registered voice data storage unit 212 .

[0111] The feature extraction unit 214 extracts a feature of the recognition target speech data acquired by the recognition target speech data acquisition unit 211. The feature is, for example, an i-vector.

[0112] The feature extraction unit 215 extracts a feature from the registration voice data acquired by the registration voice data acquisition unit 213. The feature is, for example, an i-vector.

[0113] The speaker recognition model storage unit 216 pre-stores: a first speaker recognition model machine-learned using male speech data for identifying male speakers; and a second speaker recognition model machine-learned using female speech data for identifying female speakers. The speaker recognition model storage unit 216 pre-stores the first and second speaker recognition models generated by the speaker recognition model generation device 4, described later. The methods for generating the first and second speaker recognition models will be described later.

[0114] If the gender of either the speaker of the target speech data or the speaker of the registered speech data is male, the model selection unit 217 selects a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers. Furthermore, if the gender of either the speaker of the target speech data or the speaker of the registered speech data is female, the model selection unit 217 selects a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers.

[0115] In this first embodiment, model selection unit 217 selects the first speaker recognition model when the speaker gender of the registered speech data is male, and selects the second speaker recognition model when the speaker gender of the registered speech data is female. The registered speech data is pre-associated with gender. Therefore, model selection unit 217 selects the first speaker recognition model when the gender associated with the registered speech data is male, and selects the second speaker recognition model when the gender associated with the registered speech data is female.

[0116] The speaker recognition unit 218 recognizes the speaker of the recognition target speech data by inputting the feature amount of the recognition target speech data and the feature amount of the registration speech data into either the first speaker recognition model or the second speaker recognition model selected by the model selection unit 217 .

[0117] The speaker recognition unit 218 includes a similarity calculation unit 231 and a similarity determination unit 232 .

[0118] The similarity calculation unit 231 inputs the feature value of the recognition target speech data and each feature value of the multiple registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model, and obtains the similarity between the recognition target speech data and each of the multiple registration speech data from any one of the first speaker recognition model and the second speaker recognition model.

[0119] The similarity determination unit 232 recognizes the speaker of the acquired registration voice data having the highest similarity as the speaker of the recognition target voice data.

[0120] Furthermore, the similarity determination unit 232 can also determine whether the highest similarity is greater than a threshold. Even if there is no registered voice data in the registered voice data storage unit 212 whose speaker is the same as the speaker of the target voice data, the similarity between the target voice data and each registered voice data can still be calculated. Therefore, even if the registered voice data has the highest similarity, the speaker of that registered voice data may not necessarily be the same as the speaker of the target voice data. Therefore, by determining whether the highest similarity is greater than a threshold, reliable speaker identification is possible.

[0121] The recognition result output unit 219 outputs the recognition result of the speaker identification unit 218. The recognition result output unit 219 is, for example, a display or a speaker. If the speaker of the speech data to be recognized is recognized, the recognition result output unit 219 outputs a message to the display or speaker indicating that the speaker of the speech data to be recognized is a pre-registered speaker. On the other hand, if the speaker of the speech data to be recognized is not recognized, the recognition result output unit 219 outputs a message to the display or speaker indicating that the speaker of the speech data to be recognized is not a pre-registered speaker. The recognition result output unit 219 may also output the recognition result of the speaker identification unit 218 to a device other than the speaker identification apparatus 2.

[0122] Next, the gender identification model generation device according to the first embodiment of the present invention will be described.

[0123] Figure 2 This is a diagram showing the configuration of a gender identification model generation device according to the first embodiment of the present invention.

[0124] Figure 2 The illustrated gender recognition model generation device 3 includes a gender recognition speech data storage unit 301 , a gender recognition speech data acquisition unit 302 , a feature extraction unit 303 , a gender recognition model generation unit 304 , and a gender recognition model storage unit 305 .

[0125] The gender recognition speech data acquisition unit 302, the feature extraction unit 303, and the gender recognition model generation unit 304 are implemented by a processor. The gender recognition speech data storage unit 301 and the gender recognition model storage unit 305 are implemented by a memory.

[0126] The gender identification voice data storage unit 301 stores a plurality of voice data items to which gender labels indicating male or female are assigned in advance. The gender identification voice data storage unit 301 stores a plurality of voice data items that are different from each other for each of a plurality of speakers.

[0127] The gender recognition voice data acquisition unit 302 acquires a plurality of voice data items, each of which is assigned a gender label indicating male or female, from the gender recognition voice data storage unit 301. In Embodiment 1, the gender recognition voice data acquisition unit 302 acquires a plurality of voice data items from the gender recognition voice data storage unit 301. However, the present invention is not particularly limited to this embodiment, and the plurality of voice data items may also be acquired (received) from an external device via a network.

[0128] The feature extraction unit 303 extracts feature values ​​from the plurality of speech data acquired by the gender identification speech data acquisition unit 302. The feature values ​​are, for example, i-vectors.

[0129] The gender recognition model generation unit 304 uses the feature values ​​of the first and second speech data and the similarity between the gender labels of the first and second speech data as training data, and generates a gender recognition model through machine learning that takes the feature values ​​of the two speech data as input and outputs the similarity between the two speech data. For example, the gender recognition model performs machine learning so that if the gender label of the first speech data is the same as that of the second speech data, the highest similarity is output, and if the gender label of the first speech data is different from that of the second speech data, the lowest similarity is output.

[0130] A gender recognition model using Probabilistic Linear Discriminant Analysis (PLDA) is used. The PLDA model automatically selects features effective for speaker recognition from a 400-dimensional i-vector feature set and calculates the log-likelihood ratio as the similarity.

[0131] Other examples of machine learning include supervised learning, which uses teacher data that labels input information (output information) to learn the relationship between input and output; unsupervised learning, which constructs data structures based solely on unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which uses trial and error to learn behaviors that maximize rewards. Specific machine learning methods include neural networks (including deep learning using multilayer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). For machine learning in gender recognition models, any of these specific examples can be used.

[0132] The gender recognition model storage unit 305 stores the gender recognition model generated by the gender recognition model generation unit 304 .

[0133] Furthermore, the gender recognition model generation device 3 may transmit the gender recognition model stored in the gender recognition model storage unit 305 to the speaker recognition device 2. The speaker recognition device 2 may store the received gender recognition model in the gender recognition model storage unit 203. Furthermore, when the speaker recognition device 2 is manufactured, the gender recognition model generated by the gender recognition model generation device 3 may be stored in the speaker recognition device 2.

[0134] In addition, in the gender recognition model generation device 3 of this embodiment 1, a gender label indicating whether the multiple voice data are male or female is assigned, but the present invention is not particularly limited to this. Identification information used to identify the speaker can also be assigned as a label. In this case, the gender recognition model generation unit 304 uses the various feature quantities of the first voice data and the second voice data in the multiple voice data and the similarity of the various identification information of the first voice data and the second voice data as teacher data, and generates a gender recognition model through machine learning that takes the various feature quantities of the two voice data as input and the similarity of the two voice data as output. For example, the gender recognition model performs machine learning in a manner such that if the identification information of the first voice data is the same as the identification information of the second voice data, it outputs the highest similarity, and if the identification information of the first voice data is different from the identification information of the second voice data, it outputs the lowest similarity.

[0135] Next, a speaker recognition model generation device according to Embodiment 1 of the present invention will be described.

[0136] Figure 3 This is a diagram showing the configuration of a speaker recognition model generation device according to Embodiment 1 of the present invention.

[0137] Figure 3 The speaker recognition model generation device 4 shown includes a male voice data storage unit 401, a male voice data acquisition unit 402, a feature extraction unit 403, a first speaker recognition model generation unit 404, a first speaker recognition model storage unit 405, a female voice data storage unit 411, a female voice data acquisition unit 412, a feature extraction unit 413, a second speaker recognition model generation unit 414 and a second speaker recognition model storage unit 415.

[0138] The male voice data acquisition unit 402, feature extraction unit 403, first speaker recognition model generation unit 404, female voice data acquisition unit 412, feature extraction unit 413, and second speaker recognition model generation unit 414 are implemented by a processor. The male voice data storage unit 401, first speaker recognition model storage unit 405, female voice data storage unit 411, and second speaker recognition model storage unit 415 are implemented by a memory.

[0139] The male voice data storage unit 401 stores a plurality of male voice data items to which a speaker identification label for identifying a male speaker is assigned. The male voice data storage unit 401 stores a plurality of male voice data items that are different from each other for each of the plurality of speakers.

[0140] Male voice data acquisition unit 402 acquires, from male voice data storage unit 401, a plurality of male voice data sets assigned with speaker identification tags for identifying male speakers. In Embodiment 1, male voice data acquisition unit 402 acquires the plurality of male voice data sets from male voice data storage unit 401. However, the present invention is not particularly limited to this embodiment. Alternatively, the plurality of male voice data sets may be acquired (received) from an external device via a network.

[0141] The feature extraction unit 403 extracts feature values ​​from the plurality of male voice data acquired by the male voice data acquisition unit 402. The feature values ​​are, for example, i-vectors.

[0142] The first speaker recognition model generation unit 404 uses the feature values ​​of the first male voice data and the second male voice data, as well as the similarity between the speaker recognition labels of the first male voice data and the second male voice data, as training data, and generates a first speaker recognition model through machine learning. The model takes the feature values ​​of the two voice data as input and outputs the similarity between the two voice data. For example, the first speaker recognition model performs machine learning so that if the speaker recognition labels of the first male voice data and the second male voice data are the same, the model outputs the highest similarity; if the speaker recognition labels of the first male voice data and the second male voice data are different, the model outputs the lowest similarity.

[0143] The first speaker recognition model uses a model using PLDA. The PLDA model automatically selects features effective for speaker recognition from a 400-dimensional i-vector (feature quantity) and calculates the log-likelihood ratio as the similarity.

[0144] Other examples of machine learning include supervised learning, which uses teacher data that assigns labels to input information (output information) to learn the relationship between input and output; unsupervised learning, which constructs data structures based solely on unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which uses trial and error to learn behaviors that maximize rewards. Specific machine learning methods include neural networks (including deep learning using multilayer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). Any of these specific examples can be used in the machine learning of the first speaker recognition model.

[0145] The first speaker recognition model storage unit 405 stores the first speaker recognition model generated by the first speaker recognition model generation unit 404 .

[0146] The female voice data storage unit 411 stores a plurality of female voice data items to which a speaker identification label for identifying a female speaker is assigned. The female voice data storage unit 411 stores a plurality of female voice data items that are different from each other for each of the plurality of speakers.

[0147] Female voice data acquisition unit 412 acquires, from female voice data storage unit 411, a plurality of female voice data sets, each of which has been assigned a speaker identification tag for identifying a female speaker. In Embodiment 1, female voice data acquisition unit 412 acquires the plurality of female voice data sets from female voice data storage unit 411. However, the present invention is not particularly limited to this embodiment. Alternatively, the plurality of female voice data sets may be acquired (received) from an external device via a network.

[0148] The feature extraction unit 413 extracts feature values ​​from the plurality of female voice data acquired by the female voice data acquisition unit 412. The feature values ​​are, for example, i-vectors.

[0149] Second speaker recognition model generation unit 414 uses the feature values ​​of the first and second female voice data, as well as the similarity between the speaker recognition labels of the first and second female voice data, as training data. It then generates a second speaker recognition model through machine learning, taking the feature values ​​of the two voice data as input and outputting the similarity between the two voice data. For example, the second speaker recognition model performs machine learning so that if the speaker recognition labels of the first and second female voice data are the same, the highest similarity is output; if the speaker recognition labels of the first and second female voice data are different, the lowest similarity is output.

[0150] The second speaker recognition model uses a model using PLDA. The PLDA model automatically selects features effective for speaker recognition from a 400-dimensional i-vector (feature quantity) and calculates the log-likelihood ratio as the similarity.

[0151] Other examples of machine learning include supervised learning, which uses teacher data that assigns labels to input information (output information) to learn the relationship between input and output; unsupervised learning, which constructs data structures based solely on unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which uses trial and error to learn behaviors that maximize rewards. Specific machine learning methods include neural networks (including deep learning using multilayer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). Any of these specific examples can be used in the machine learning of the second speaker recognition model.

[0152] The second speaker recognition model storage unit 415 stores the second speaker recognition model generated by the second speaker recognition model generation unit 414 .

[0153] Furthermore, speaker recognition model generation device 4 may transmit the first speaker recognition model stored in first speaker recognition model storage unit 405 and the second speaker recognition model stored in second speaker recognition model storage unit 415 to speaker recognition device 2. Speaker recognition device 2 may store the received first speaker recognition model and second speaker recognition model in speaker recognition model storage unit 216. Furthermore, when speaker recognition device 2 is manufactured, the first speaker recognition model and second speaker recognition model generated by speaker recognition model generation device 4 may be stored in speaker recognition device 2.

[0154] Next, the operations of the registration process and the speaker recognition process of the speaker recognition device 2 according to the first embodiment will be described.

[0155] Figure 4 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the first embodiment.

[0156] First, in step S1, the registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1. A speaker who wishes to register their voice data speaks a predetermined passage into the microphone 1. In this case, the passage in the registration target voice data is preferably longer than the passage in the recognition target voice data. Acquiring registration target voice data with a relatively large number of characters can improve speaker recognition accuracy. Alternatively, the speaker recognition device 2 may present the registration target speaker with multiple predetermined passages. In this case, the registration target speaker speaks the multiple passages presented.

[0157] Next, in step S2 , the feature amount extraction unit 202 extracts the feature amount of the registration target speech data acquired by the registration target speech data acquisition unit 201 .

[0158] Next, in step S3, the gender recognition unit 205 performs a gender recognition process for recognizing the gender of the speaker of the registration target voice data using the feature quantity of the registration target voice data extracted by the feature quantity extraction unit 202. The gender recognition process will be described later.

[0159] Next, in step S4 , the registration unit 206 stores the registration target voice data corresponding to the gender information identified by the gender identification unit 205 in the registration voice data storage unit 212 as registration voice data.

[0160] Figure 5 Is used to illustrate Figure 4 Flowchart of the operation of the gender recognition processing of step S3.

[0161] First, in step S11, the gender recognition unit 205 obtains a gender recognition model from the gender recognition model storage unit 203. The gender recognition model is a gender recognition model that is machine-learned using the feature values ​​of the first and second speech data, as well as the similarity between the gender labels of the first and second speech data, as training data. Alternatively, the gender recognition model may be a gender recognition model that is machine-learned using the feature values ​​of the first and second speech data, as well as the similarity between the identification information of the first and second speech data, as training data.

[0162] Next, in step S12 , the gender recognition unit 205 acquires the feature amount of the male voice data from the gender recognition voice data storage unit 204 .

[0163] Next, in step S13 , the gender recognition section 205 calculates the similarity between the registration object voice data and the male voice data by inputting the feature amount of the registration object voice data and the feature amount of the male voice data into the gender recognition model.

[0164] Next, in step S14, the gender recognition unit 205 determines whether similarities have been calculated between the target voice data and all male voice data. If similarities have not been calculated between the target voice data and all male voice data (No in step S14), the process returns to step S12. The gender recognition unit 205 then obtains, from the gender recognition voice data storage unit 204, the feature values ​​of the male voice data for which similarities have not been calculated, from the feature values ​​of the plurality of male voice data.

[0165] On the other hand, if it is determined that the similarities between the registration target voice data and all male voice data have been calculated (Yes in step S14 ), in step S15 , the gender identification unit 205 calculates the average of the calculated similarities as the average male similarity.

[0166] Next, in step S16 , the gender recognition unit 205 acquires the feature value of the female voice data from the gender recognition voice data storage unit 204 .

[0167] Next, in step S17 , the gender recognition section 205 calculates the similarity between the registration target voice data and the female voice data by inputting the feature amount of the registration target voice data and the feature amount of the female voice data into the gender recognition model.

[0168] Next, in step S18, the gender recognition unit 205 determines whether similarities have been calculated between the target voice data and all female voice data. If it is determined that similarities have not been calculated between the target voice data and all female voice data (No in step S18), the process returns to step S16. The gender recognition unit 205 then obtains, from the gender recognition voice data storage unit 204, the feature values ​​of the female voice data for which similarities have not been calculated, from the feature values ​​of the plurality of female voice data.

[0169] On the other hand, if it is determined that the similarities between the registration target voice data and all female voice data have been calculated (Yes in step S18 ), in step S19 , the gender identification unit 205 calculates the average of the calculated similarities as the average female similarity.

[0170] Next, in step S20 , the gender recognition unit 205 outputs the higher gender between the average male similarity and the average female similarity as a recognition result to the registration unit 206 .

[0171] Figure 6 is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the first embodiment. Figure 7 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition apparatus according to the first embodiment.

[0172] First, in step S31, the recognition target speech data acquisition unit 211 acquires the recognition target speech data output from the microphone 1. The recognition target speaker speaks to the microphone 1. The microphone 1 picks up the voice of the recognition target speaker and outputs the recognition target speech data.

[0173] Next, in step S32 , the feature amount extraction unit 214 extracts the feature amount of the recognition target speech data acquired by the recognition target speech data acquisition unit 211 .

[0174] Next, in step S33 , the registered voice data acquisition unit 213 acquires the registered voice data from the registered voice data storage unit 212 . At this time, the registered voice data acquisition unit 213 acquires one piece of registered voice data from the plurality of registered voice data stored in the registered voice data storage unit 212 .

[0175] Next, in step S34 , the feature amount extraction unit 215 extracts the feature amount of the registration voice data acquired by the registration voice data acquisition unit 213 .

[0176] Next, in step S35 , the model selection unit 217 acquires the gender corresponding to the registration voice data acquired by the registration voice data acquisition unit 213 .

[0177] Next, in step S36, the model selection unit 217 determines whether the acquired gender is male. If the acquired gender is male (yes in step S36), the model selection unit 217 selects the first speaker identification model in step S37. The model selection unit 217 retrieves the selected first speaker identification model from the speaker identification model storage unit 216 and outputs the retrieved first speaker identification model to the similarity calculation unit 231.

[0178] On the other hand, if the acquired gender is determined not to be male, that is, if the acquired gender is determined to be female (No in step S36), in step S38, the model selection unit 217 selects the second speaker identification model. The model selection unit 217 obtains the selected second speaker identification model from the speaker identification model storage unit 216 and outputs the obtained second speaker identification model to the similarity calculation unit 231.

[0179] Next, in step S39 , the similarity calculation unit 231 calculates the similarity between the recognition target speech data and the registration speech data by inputting the feature quantity of the recognition target speech data and the feature quantity of the registration speech data into any one of the selected first speaker recognition model and second speaker recognition model.

[0180] Next, in step S40, the similarity calculation unit 231 determines whether similarities have been calculated between the target speech data and all registered speech data stored in the registered speech data storage unit 212. If it is determined that similarities have not been calculated between the target speech data and all registered speech data (No in step S40), the process returns to step S33. The registered speech data acquisition unit 213 then acquires, from the plurality of registered speech data stored in the registered speech data storage unit 212, any registered speech data for which similarities have not been calculated.

[0181] On the other hand, when it is determined that the similarities between the recognition target voice data and all the registered voice data have been calculated (YES in step S40 ), the similarity determination unit 232 determines whether the highest similarity is greater than a threshold value in step S41 .

[0182] Here, when it is determined that the highest similarity is greater than the threshold (Yes in step S41 ), in step S42 , the similarity determination unit 232 recognizes the speaker of the registration voice data with the highest similarity as the speaker of the recognition target voice data.

[0183] On the other hand, if the highest similarity is determined to be below the threshold (No in step S41 ), the similarity determination unit 232 determines in step S43 that there is no registered voice data having the same speaker as the recognition target voice data among the plurality of registered voice data.

[0184] Next, in step S44, the recognition result output unit 219 outputs the recognition result of the speaker identification unit 218. If the recognition result output unit 219 identifies the speaker of the speech data to be recognized, it outputs a message indicating that the speaker of the speech data to be recognized is a pre-registered speaker. On the other hand, if the recognition result output unit 219 does not identify the speaker of the speech data to be recognized, it outputs a message indicating that the speaker of the speech data to be recognized is not a pre-registered speaker.

[0185] As described above, when either the gender of the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is recognized by inputting the feature quantity of the speech data to be recognized and the feature quantity of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, when either the gender of the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is recognized by inputting the feature quantity of the speech data to be recognized and the feature quantity of the registered speech data into a second speaker recognition model generated for female identification.

[0186] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model and the second speaker recognition model specifically for each gender, thereby improving the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker.

[0187] Furthermore, in this first embodiment, the model selection unit 217 selects either the first or second speaker recognition model based on the gender of the speaker in the enrollment speech data, but the present invention is not particularly limited to this. The model selection unit 217 may also select either the first or second speaker recognition model based on the gender of the speaker in the speech data to be recognized. In this case, the speaker recognition device 2 includes: a gender recognition unit that recognizes the gender of the speaker in the speech data to be recognized; a gender recognition model storage unit that pre-stores a gender recognition model machine-learned using male and female speech data for speaker gender recognition; and a gender recognition speech data storage unit that pre-stores feature values ​​of the gender recognition speech data used to recognize the gender of the speaker in the speech data to be recognized. The configurations of the gender recognition unit, gender recognition model storage unit, and gender recognition speech data storage unit are similar to those of the gender recognition unit 205, gender recognition model storage unit 203, and gender recognition speech data storage unit 204 described above. Furthermore, when the gender of the speaker of the recognition target speech data can be identified, the gender recognition unit 205 , the gender recognition model storage unit 203 , and the gender recognition speech data storage unit 204 are unnecessary.

[0188] Next, the operation of the gender identification model generation process performed by the gender identification model generation device 3 according to the first embodiment will be described.

[0189] Figure 8 This is a flowchart for explaining the operation of the gender identification model generation process performed by the gender identification model generation device according to the first embodiment.

[0190] First, in step S51 , the gender identification voice data acquisition unit 302 acquires a plurality of voice data items to which a gender label indicating male or female is assigned from the gender identification voice data storage unit 301 .

[0191] Next, in step S52 , the feature amount extraction unit 303 extracts feature amounts of the plurality of speech data acquired by the gender identification speech data acquisition unit 302 .

[0192] Next, in step S53 , the gender recognition model generation unit 304 obtains, as training data, the feature values ​​of the first and second speech data and the similarity between the gender labels of the first and second speech data from the plurality of speech data.

[0193] Next, in step S54, the gender recognition model generation unit 304 uses the acquired teacher data to machine-learn a gender recognition model that takes the feature quantities of the two speech data as input and takes the similarity of the two speech data as output.

[0194] Next, in step S55, the gender recognition model generation unit 304 determines whether a gender recognition model has been learned using the combination of all the speech data. If it is determined that a gender recognition model has not been learned using the combination of all the speech data (No in step S55), the process returns to step S53. The gender recognition model generation unit 304 then obtains, as training data, the feature values ​​of the first and second speech data sets that were not used for machine learning, as well as the similarity of the gender labels of the first and second speech data sets.

[0195] On the other hand, when it is determined that the gender recognition model has been learned using the combined machine learning of all speech data (yes in step S55 ), in step S56 , the gender recognition model generation unit 304 stores the gender recognition model generated by machine learning in the gender recognition model storage unit 305 .

[0196] As described above, by inputting the feature values ​​of the registration voice data or the recognition target voice data and the feature values ​​of the male voice data into the gender recognition model generated by machine learning, a first similarity between the two voice data is output. Furthermore, by inputting the feature values ​​of the registration voice data or the recognition target voice data and the feature values ​​of the female voice data into the gender recognition model, a second similarity between the two voice data is output. Then, by comparing the first and second similarities, the gender of the speaker of the registration voice data or the recognition target voice data can be easily estimated.

[0197] Next, the operation of speaker recognition model generation processing by speaker recognition model generation device 4 according to the first embodiment will be described.

[0198] Figure 9 This is a first flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the first embodiment. Figure 10 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the first embodiment.

[0199] First, in step S61 , the male voice data acquisition unit 402 acquires a plurality of male voice data items to which speaker identification labels for identifying male speakers are assigned from the male voice data storage unit 401 .

[0200] Next, in step S62 , the feature amount extraction unit 403 extracts feature amounts of the plurality of male voice data acquired by the male voice data acquisition unit 402 .

[0201] Next, in step S63 , the first speaker recognition model generation unit 404 obtains the feature values ​​of the first male voice data and the second male voice data and the similarity of the speaker recognition labels of the first male voice data and the second male voice data from the plurality of male voice data as teacher data.

[0202] Next, in step S64 , the first speaker recognition model generation unit 404 uses the acquired teacher data to machine-learn a first speaker recognition model that takes the feature quantities of the two speech data as input and outputs the similarity between the two speech data.

[0203] Next, in step S65, the first speaker recognition model generation unit 404 determines whether the first speaker recognition model has been trained using the combination of all male speech data from the plurality of male speech data. If it is determined that the first speaker recognition model has not been trained using the combination of all male speech data (No in step S65), the process returns to step S63. The first speaker recognition model generation unit 404 then obtains, as training data, the feature values ​​of the first and second male speech data from the plurality of male speech data that were not used for machine learning, as well as the similarity between the speaker recognition labels of the first and second male speech data.

[0204] On the other hand, when it is determined that the first speaker recognition model has been machine-learned using the combination of all male voice data (yes in step S65), in step S66, the first speaker recognition model generation unit 404 stores the first speaker recognition model generated by machine learning in the first speaker recognition model storage unit 405.

[0205] Next, in step S67 , the female voice data acquisition unit 412 acquires a plurality of female voice data items to which speaker identification labels for identifying female speakers are assigned from the female voice data storage unit 411 .

[0206] Next, in step S68 , the feature amount extraction unit 413 extracts feature amounts of the plurality of female voice data acquired by the female voice data acquisition unit 412 .

[0207] Next, in step S69, the second speaker recognition model generation unit 414 obtains the feature values ​​of the first female voice data and the second female voice data and the similarity of the speaker recognition labels of the first female voice data and the second female voice data from the plurality of female voice data as teacher data.

[0208] Next, in step S70 , the second speaker recognition model generation unit 414 uses the acquired teacher data to machine-learn a second speaker recognition model that takes the feature quantities of the two speech data as input and outputs the similarity between the two speech data.

[0209] Next, in step S71, the second speaker recognition model generation unit 414 determines whether the second speaker recognition model has been trained using the combination of all female voice data from the plurality of female voice data. If it is determined that the second speaker recognition model has not been trained using the combination of all female voice data (No in step S71), the process returns to step S69. The second speaker recognition model generation unit 414 then obtains, as training data, the feature values ​​of the first and second female voice data from the plurality of female voice data that were not used for machine learning, as well as the similarity between the speaker recognition labels of the first and second female voice data.

[0210] On the other hand, when it is determined that the second speaker recognition model has been machine-learned using the combination of all female voice data (yes in step S71), in step S72, the second speaker recognition model generation unit 414 stores the second speaker recognition model generated by machine learning in the second speaker recognition model storage unit 415.

[0211] As described above, when either the gender of the speaker of the speech data to be recognized or the speaker of the registered speech data is male, the speaker of the speech data to be recognized is recognized by inputting the feature quantity of the speech data to be recognized and the feature quantity of the registered speech data into a first speaker recognition model generated for male identification. Furthermore, when either the gender of the speaker of the speech data to be recognized or the speaker of the registered speech data is female, the speaker of the speech data to be recognized is recognized by inputting the feature quantity of the speech data to be recognized and the feature quantity of the registered speech data into a second speaker recognition model generated for female identification.

[0212] Therefore, even if the feature quantity distribution of speech data differs according to gender, the speaker of the speech data of the recognition object is identified using the first speaker recognition model specifically for males and the second speaker recognition model specifically for females. Therefore, the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker can be improved.

[0213] Next, evaluation of the speaker recognition performance of the speaker recognition apparatus 2 according to the first embodiment will be described.

[0214] Figure 111 is a diagram showing speaker recognition performance evaluation results of a conventional speaker recognition apparatus and speaker recognition performance evaluation results of the speaker recognition apparatus according to Embodiment 1. FIG.

[0215] Figure 11 The performance evaluation results shown are the results of speaker recognition using conventional speaker recognition apparatuses and the speaker recognition apparatus 2 of the first embodiment for speech data provided by the SRE19 progress dataset and the SRE19 evaluation dataset.

[0216] SRE19 is a speaker recognition competition hosted by the National Institute of Standards and Technology (NIST). Both the SRE19 Progress Dataset and the SRE19 Evaluation Dataset are provided by SRE19.

[0217] Conventional speaker recognition devices recognize the speaker of speech data using a speaker recognition model generated without distinguishing between males and females.

[0218] Furthermore, the speaker identification apparatus 2 of the first embodiment identifies the speaker of the speech data using the first speaker identification model for males and the second speaker identification model for females.

[0219] Evaluation results are expressed as EER (Equal Error Rate) (%), a metric commonly used for speaker recognition evaluation, minC (minimum cost), and actC (actual cost), also defined by NIST. Smaller EER, minC, and actC values ​​indicate higher performance.

[0220] like Figure 11 As shown in FIG1 , the EER, minC, and actC of the speaker recognition apparatus 2 of the first embodiment are all smaller than those of conventional speaker recognition apparatuses. Therefore, it can be seen that the speaker recognition performance of the speaker recognition apparatus 2 of the first embodiment is higher than that of conventional speaker recognition apparatuses.

[0221] Next, a speaker identification device 2 according to a first modification of the first embodiment will be described.

[0222] The gender identification process performed by the gender identification unit 205 of the speaker identification apparatus 2 according to the first modification of the first embodiment is different from the above-described gender identification process.

[0223] In the first variation of the first embodiment, the gender recognition unit 205 inputs the feature value of the registration target voice data and the feature value of each of the plurality of male voice data previously stored in the gender recognition voice data storage unit 204 into the gender recognition model. The unit then obtains the similarity between the registration target voice data and each of the plurality of male voice data from the gender recognition model. The gender recognition unit 205 then calculates the maximum of the plurality of similarities obtained as the maximum male similarity.

[0224] Furthermore, the gender recognition unit 205 inputs the feature values ​​of the registration target voice data and each of the plurality of female voice data previously stored in the gender recognition voice data storage unit 204 into the gender recognition model. The gender recognition unit 205 then obtains the similarity between the registration target voice data and each of the plurality of female voice data from the gender recognition model. The gender recognition unit 205 then calculates the maximum of the plurality of similarities obtained as the maximum female similarity.

[0225] If the maximum male similarity is higher than the maximum female similarity, the gender recognition unit 205 recognizes the gender of the speaker of the target voice data as male. On the other hand, if the maximum male similarity is lower than the maximum female similarity, the gender recognition unit 205 recognizes the gender of the speaker of the target voice data as female. Furthermore, if the maximum male similarity is equal to the maximum female similarity, the gender recognition unit 205 may recognize the gender of the speaker of the target voice data as male or female.

[0226] Next, the operation of the gender identification process according to the first modification of the first embodiment will be described.

[0227] Figure 12 This is a flowchart for explaining the operation of the gender recognition process of the first modification of the present embodiment 1. The gender recognition process of the first modification of the present embodiment 1 is Figure 4 Another example of the gender recognition process in step S3.

[0228] The processing of steps S81 to S84 is Figure 5 The processes of steps S11 to S14 shown are the same, and therefore their description is omitted.

[0229] If it is determined that the similarities between the registration target voice data and all male voice data have been calculated (YES in step S84 ), in step S85 , the gender identification unit 205 calculates the maximum value of the calculated similarities as the maximum male similarity.

[0230] The processing of steps S86 to S88 is Figure 5The processes of steps S16 to S18 are the same, so their description is omitted.

[0231] If it is determined that the similarities between the registration target voice data and all female voice data have been calculated (YES in step S88 ), in step S89 , the gender identification unit 205 calculates the maximum value among the calculated similarities as the maximum female similarity.

[0232] Next, in step S90 , the gender recognition unit 205 outputs the higher gender between the maximum male similarity and the maximum female similarity as the recognition result to the registration unit 206 .

[0233] Next, a speaker identification device 2 according to a second modification of the first embodiment will be described.

[0234] The gender identification process performed by the gender identification unit 205 of the speaker identification apparatus 2 according to the second modification of the first embodiment is different from the above-described gender identification process.

[0235] In the second variation of the first embodiment, the gender recognition unit 205 calculates the average feature value of a plurality of pre-stored male speech data. Furthermore, the gender recognition unit 205 inputs the feature value of the registration target speech data and the average feature value of the plurality of male speech data into a gender recognition model, and obtains a first similarity between the registration target speech data and the male speech data group from the gender recognition model.

[0236] Furthermore, the gender recognition unit 205 calculates the average feature value of a plurality of pre-stored female voice data. Furthermore, the gender recognition unit 205 inputs the feature value of the registration target voice data and the average feature value of the plurality of female voice data into a gender recognition model, and obtains a second similarity between the registration target voice data and the female voice data group from the gender recognition model.

[0237] When the first similarity is higher than the second similarity, the gender recognition unit 205 recognizes the gender of the speaker of the target voice data as male. On the other hand, when the first similarity is lower than the second similarity, the gender recognition unit 205 recognizes the gender of the speaker of the target voice data as female. Furthermore, when the first similarity and the second similarity are equal, the gender recognition unit 205 may recognize the gender of the speaker of the target voice data as male or female.

[0238] Next, the operation of the gender identification process according to the second modification of the first embodiment will be described.

[0239] Figure 13This is a flowchart for explaining the operation of the gender recognition process of the second modification of the first embodiment. Figure 4 Another example of the gender recognition process in step S3.

[0240] First, in step S101, the gender recognition unit 205 obtains a gender recognition model from the gender recognition model storage unit 203. The gender recognition model is a gender recognition model that is machine-learned using the feature values ​​of the first and second speech data, as well as the similarity between the gender labels of the first and second speech data, as training data. Alternatively, the gender recognition model may be a gender recognition model that is machine-learned using the feature values ​​of the first and second speech data, as well as the similarity between the identification information of the first and second speech data, as training data.

[0241] Next, in step S102 , the gender recognition unit 205 acquires feature quantities of a plurality of male voice data from the gender recognition voice data storage unit 204 .

[0242] Next, in step S103 , the gender identification unit 205 calculates the average feature value of the acquired plurality of male voice data.

[0243] Next, in step S104 , the gender recognition unit 205 calculates a first similarity between the registration target voice data and the male voice data group by inputting the feature value of the registration target voice data and the average feature value of a plurality of male voice data into the gender recognition model.

[0244] Next, in step S105 , the gender recognition unit 205 acquires feature quantities of a plurality of female voice data from the gender recognition voice data storage unit 204 .

[0245] Next, in step S106 , the gender identification unit 205 calculates the average feature value of the acquired plurality of female voice data.

[0246] Next, in step S107 , the gender identification unit 205 calculates a second similarity between the registration target voice data and the female voice data group by inputting the feature value of the registration target voice data and the average feature value of a plurality of female voice data into the gender identification model.

[0247] Next, in step S108 , the gender recognition unit 205 outputs the gender having the higher of the first similarity and the second similarity as a recognition result to the registration unit 206 .

[0248] In addition, in the gender recognition processing of this first embodiment, the similarity between the registered voice data and male voice data, as well as the similarity between the registered voice data and female voice data, is calculated, and the two calculated similarities are compared to determine the gender of the registered voice data, but the present invention is not particularly limited to this. For example, the gender recognition model generation unit 304 may also use the feature value of one voice data and the gender label of one voice data among multiple voice data as training data, and generate a gender recognition model through machine learning that takes the feature value of the voice data as input and outputs either male or female. In this case, the gender recognition model is, for example, a deep neural network model, and the machine learning is, for example, deep learning.

[0249] Alternatively, the gender recognition model may calculate male and female probabilities for input speech data. In this case, the gender recognition unit 205 may output the gender with the higher male and female probabilities as the recognition result.

[0250] Furthermore, the speaker recognition device 2 may further include an input receiving unit that receives input of the gender of the speaker of the target voice data when acquiring the target voice data. In this case, the registration unit 206 may store the target voice data acquired by the target voice data acquisition unit 201 and the gender received by the input receiving unit in the registration voice data storage unit 212, in association with each other. This eliminates the need for the feature extraction unit 202, gender recognition model storage unit 203, gender recognition voice data storage unit 204, and gender recognition unit 205, simplifying the structure of the speaker recognition device 2 and reducing the load on the registration process of the speaker recognition device 2.

[0251] (Implementation Method 2)

[0252] In the first embodiment described above, the similarity between each of the registered voice data stored in the registered voice data storage unit 212 and the target voice data is calculated, and the speaker of the registered voice data with the highest similarity is identified as the speaker of the target voice data. In contrast, in the second embodiment, the identification information of the speaker of the target voice data is input, and one piece of registered voice data, pre-associated with the identification information, is retrieved from the plurality of registered voice data stored in the registered voice data storage unit 212. The similarity between the one piece of registered voice data and the target voice data is then calculated. If the similarity exceeds a threshold, the speaker of the registered voice data is identified as the speaker of the target voice data.

[0253] Figure 14 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 2 of the present invention.

[0254] Figure 14 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 21. The speaker recognition device 21 may or may not include the microphone 1.

[0255] In the second embodiment, the same components as those in the first embodiment are denoted by the same reference numerals and their descriptions are omitted.

[0256] The speaker recognition device 21 includes a registration object voice data acquisition unit 201, a feature value extraction unit 202, a gender recognition model storage unit 203, a gender recognition voice data storage unit 204, a gender recognition unit 205, a registration unit 2061, a recognition object voice data acquisition unit 211, a registration voice data storage unit 2121, a registration voice data acquisition unit 2131, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 216, a model selection unit 217, a speaker recognition unit 2181, a recognition result output unit 219, an input acceptance unit 221 and a recognition information acquisition unit 222.

[0257] The input receiving unit 221 is, for example, an input device such as a keyboard, a mouse, or a touch screen. When registering voice data, the input receiving unit 221 receives identification information input by the speaker for identifying the speaker of the voice data being registered. Furthermore, when registering voice data, the input receiving unit 221 receives identification information input by the speaker for identifying the speaker of the voice data to be recognized. Alternatively, the input receiving unit 221 may be a card reader, an RFID (Radio Frequency Identification) reader, or the like. In this case, the speaker uses a card reader to read a card containing identification information, or uses an RFID reader to read an RFID tag containing identification information.

[0258] The identification information acquisition unit 222 acquires the identification information received by the input acceptance unit 221. When registering voice data, the identification information acquisition unit 222 acquires identification information for identifying the speaker of the voice data to be registered and transmits the acquired identification information to the registration unit 2061. Furthermore, when recognizing voice data, the identification information acquisition unit 222 acquires identification information for identifying the speaker of the voice data to be recognized and transmits the acquired identification information to the registration voice data acquisition unit 2131.

[0259] The registration unit 2061 registers, as registration voice data, the registration target voice data that is associated with the gender information identified by the gender identification unit 205 and the identification information acquired by the identification information acquisition unit 222. The registration unit 2061 registers the registration voice data in the registration voice data storage unit 2121.

[0260] The registration voice data storage unit 2121 stores registration voice data in association with gender information and identification information. The registration voice data storage unit 2121 stores a plurality of registration voice data. The plurality of registration voice data is associated with identification information for identifying the speaker of each of the plurality of registration voice data.

[0261] The registered voice data acquisition unit 2131 acquires the registered voice data corresponding to the identification information that matches the identification information acquired by the identification information acquisition unit 222 from the plurality of registered voice data registered in the registered voice data storage unit 2121 .

[0262] The speaker recognition unit 2181 includes a similarity calculation unit 2311 and a similarity determination unit 2321 .

[0263] The similarity calculation unit 2311 obtains the similarity between the recognition object speech data and the registration speech data from any one of the first speaker recognition model and the second speaker recognition model by inputting the feature value of the recognition object speech data and the feature value of the registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model.

[0264] When the acquired similarity is higher than the threshold, the similarity determination unit 2321 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0265] Next, the operations of the registration process and the speaker recognition process of the speaker recognition device 21 according to the second embodiment will be described.

[0266] Figure 15 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the second embodiment of the present invention.

[0267] First, in step S121, the identification information acquisition unit 222 acquires the speaker's identification information received by the input acceptance unit 221. The input acceptance unit 221 accepts the speaker's identification information input for identifying the registered voice data and outputs the received identification information to the identification information acquisition unit 222. The identification information acquisition unit 222 outputs the identification information for identifying the speaker of the voice data to be registered to the registration unit 2061.

[0268] Next, in step S122, the registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1. A speaker who inputs recognition information and wishes to register his / her own voice data speaks a predetermined text into the microphone 1.

[0269] The processing of step S123 and step S124 is the same as Figure 4 The processes of step S2 and step S3 shown are the same, and therefore their description is omitted.

[0270] Next, in step S125, the registration unit 2061 stores the registration target voice data, which is associated with the gender information identified by the gender recognition unit 205 and the identification information acquired by the identification information acquisition unit 222, as registration voice data in the registration voice data storage unit 2121. Consequently, the registration voice data storage unit 2121 stores the registration voice data associated with the gender information and the identification information.

[0271] Figure 16 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the second embodiment. Figure 17 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the second embodiment.

[0272] First, in step S131, the identification information acquisition unit 222 acquires the identification information of the speaker received by the input acceptance unit 221. The input acceptance unit 221 accepts the identification information input by the speaker for identifying the speaker of the speech data to be recognized and outputs the received identification information to the identification information acquisition unit 222. The identification information acquisition unit 222 outputs the identification information used to identify the speaker of the recognition target speech data to the registered speech data acquisition unit 2131.

[0273] Next, in step S132, the registered voice data acquisition unit 2131 determines whether the registered voice data storage unit 2121 contains identification information that matches the identification information acquired by the identification information acquisition unit 222. If the registered voice data storage unit 2121 determines that the acquired identification information does not match the acquired identification information (No in step S132), the speaker recognition process ends. Alternatively, if the registered voice data storage unit 2121 determines that the acquired identification information does not match the acquired identification information, the recognition result output unit 219 may output notification information to notify the speaker that the input identification information has not been registered. Furthermore, if the registered voice data storage unit 2121 determines that the acquired identification information does not match the acquired identification information, the recognition result output unit 219 may output notification information to prompt the speaker to register voice data.

[0274] On the other hand, when it is determined that the registered voice data storage unit 2121 contains identification information that matches the acquired identification information (YES in step S132 ), in step S133 , the recognition target voice data acquisition unit 211 acquires the recognition target voice data output from the microphone 1 .

[0275] Next, in step S134 , the feature amount extraction unit 214 extracts the feature amount of the recognition target speech data acquired by the recognition target speech data acquisition unit 211 .

[0276] Next, in step S135 , the registered voice data acquisition unit 2131 acquires the registered voice data corresponding to the identification information that matches the identification information acquired by the identification information acquisition unit 222 from the plurality of registered voice data registered in the registered voice data storage unit 2121 .

[0277] The processing of steps S136 to S141 is Figure 6 The processes of steps S34 to S39 shown are the same, so their description is omitted.

[0278] Next, in step S142 , the similarity determination unit 2321 determines whether the similarity calculated by the similarity calculation unit 2311 is greater than a threshold value.

[0279] Here, when it is determined that the similarity calculated by the similarity calculation unit 2311 is greater than the threshold (Yes in step S142 ), in step S143 , the similarity determination unit 2321 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0280] On the other hand, when the similarity calculation unit 2311 determines that the similarity calculated is below the threshold (No in step S142), the similarity determination unit 2321 determines in step S144 that the speaker of the recognition target speech data is not the speaker of the registered speech data.

[0281] Next, in step S145, the recognition result output unit 219 outputs the recognition result of the speaker identification unit 2181. If the recognition result output unit 219 identifies the speaker of the speech data to be recognized, it outputs a message indicating that the speaker of the speech data to be recognized is a pre-registered speaker. On the other hand, if the recognition result output unit 219 does not identify the speaker of the speech data to be recognized, it outputs a message indicating that the speaker of the speech data to be recognized is not a pre-registered speaker.

[0282] As described above, in this second embodiment, only the similarity between the registration voice data corresponding to the recognition information and the recognition target voice data is calculated. Therefore, compared to the first embodiment, which calculates multiple similarities between each of multiple registration voice data and the recognition target voice data, the processing load for calculating similarities can be reduced in the second embodiment.

[0283] (Implementation Method 3)

[0284] In the aforementioned embodiments 1 and 2, either the first or second speaker recognition model is selected based on the gender of the speaker in the enrollment speech data. However, when using two different first and second speaker recognition models, the output value ranges of the first and second speaker recognition models may differ. To address this issue, in embodiment 3, when enrolling speech data, a first threshold and a second threshold for identifying the same speaker are calculated for each of the first and second speaker recognition models, and the calculated first and second thresholds are stored. Furthermore, when recognizing speech data, the similarity between the target speech data and the enrollment speech data is corrected by subtracting the first or second threshold from the calculated similarity. The corrected similarity is then compared with a third threshold shared by both the first and second speaker recognition models to identify the speaker of the target speech data.

[0285] First, a speaker recognition model generation device according to a third embodiment of the present invention will be described.

[0286] Figure 18 This is a diagram showing the configuration of a speaker recognition model generation device according to a third embodiment of the present invention.

[0287] Figure 18 The speaker recognition model generation device 41 shown includes a male voice data storage unit 401, a male voice data acquisition unit 402, a feature extraction unit 403, a first speaker recognition model generation unit 404, a first speaker recognition model storage unit 405, a first speaker recognition unit 406, a first threshold calculation unit 407, a threshold storage unit 408, a female voice data storage unit 411, a female voice data acquisition unit 412, a feature extraction unit 413, a second speaker recognition model generation unit 414, a second speaker recognition model storage unit 415, a second speaker recognition unit 416 and a second threshold calculation unit 417.

[0288] In the third embodiment, the same components as those in the first and second embodiments are denoted by the same reference numerals and their descriptions are omitted.

[0289] The first speaker recognition unit 406 inputs all combinations of feature quantities of two speech data among the plurality of male speech data into the first speaker recognition model, and obtains similarities of each of the plurality of combinations of the two speech data from the first speaker recognition model.

[0290] The first threshold calculation unit 407 calculates a first threshold that can distinguish the similarity between two speech data of the same speaker and the similarity between two speech data of different speakers. The first threshold calculation unit 407 calculates the first threshold by performing regression analysis on the multiple similarities calculated by the first speaker recognition unit 406.

[0291] The second speaker recognition unit 416 inputs all combinations of feature quantities of two speech data among the plurality of female speech data into the second speaker recognition model, and obtains similarities of each of the plurality of combinations of the two speech data from the second speaker recognition model.

[0292] The second threshold calculation unit 417 calculates a second threshold that can distinguish the similarity between two speech data of the same speaker and the similarity between two speech data of different speakers. The second threshold calculation unit 417 calculates the second threshold by performing regression analysis on the multiple similarities calculated by the second speaker recognition unit 416.

[0293] The threshold storage unit 408 stores the first threshold calculated by the first threshold calculation unit 407 and the second threshold calculated by the second threshold calculation unit 417 .

[0294] Next, a speaker recognition system according to a third embodiment of the present invention will be described.

[0295] Figure 19 This is a diagram showing the configuration of a speaker recognition system according to Embodiment 3 of the present invention.

[0296] Figure 19 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 22. The speaker recognition device 22 may or may not include the microphone 1.

[0297] In the third embodiment, the same components as those in the first and second embodiments are denoted by the same reference numerals and their descriptions are omitted.

[0298] The speaker recognition device 22 includes a registration object voice data acquisition unit 201, a feature value extraction unit 202, a gender recognition model storage unit 203, a gender recognition voice data storage unit 204, a gender recognition unit 205, a registration unit 2061, a recognition object voice data acquisition unit 211, a registration voice data storage unit 2121, a registration voice data acquisition unit 2131, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 216, a model selection unit 217, a speaker recognition unit 2182, a recognition result output unit 219, an input acceptance unit 221, a recognition information acquisition unit 222, and a threshold storage unit 223.

[0299] The speaker recognition unit 2182 includes a similarity calculation unit 2311 , a similarity correction unit 233 , and a similarity determination unit 2322 .

[0300] When the similarity is obtained from the first speaker recognition model, the similarity correction unit 233 subtracts the first threshold from the obtained similarity. Furthermore, when the similarity is obtained from the second speaker recognition model, the similarity correction unit 233 subtracts the second threshold from the obtained similarity. When the similarity calculation unit 2311 calculates similarity using the first speaker recognition model, the similarity correction unit 233 reads the first threshold from the threshold storage unit 223 and subtracts the first threshold from the calculated similarity. Furthermore, when the similarity calculation unit 2311 calculates similarity using the second speaker recognition model, the similarity correction unit 233 reads the second threshold from the threshold storage unit 223 and subtracts the second threshold from the calculated similarity.

[0301] The threshold storage unit 223 pre-stores a first threshold for correcting the similarity calculated using the first speaker recognition model and a second threshold for correcting the similarity calculated using the second speaker recognition model. The threshold storage unit 223 pre-stores the first and second thresholds generated by the speaker recognition model generation device 41.

[0302] Furthermore, the speaker recognition model generation device 41 may transmit the first and second threshold values ​​stored in the threshold value storage unit 408 to the speaker recognition device 22. The speaker recognition device 22 may store the received first and second threshold values ​​in the threshold value storage unit 223. Furthermore, when the speaker recognition device 22 is manufactured, the first and second threshold values ​​generated by the speaker recognition model generation device 41 may be stored in the threshold value storage unit 223.

[0303] When the similarity obtained by subtracting the first threshold or the second threshold by the similarity correction unit 233 exceeds the third threshold, the similarity determination unit 2322 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0304] Next, the operation of speaker recognition model generation processing by speaker recognition model generation device 41 according to the third embodiment will be described.

[0305] Figure 20 This is a first flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the third embodiment. Figure 21 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the third embodiment. Figure 22 This is a third flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the third embodiment.

[0306] The processing of steps S151 to S156 is Figure 9 The processes of steps S61 to S66 shown are the same, so their description is omitted.

[0307] Next, in step S157 , the first speaker recognition unit 406 acquires the first speaker recognition model from the first speaker recognition model storage unit 405 .

[0308] Next, in step S158 , the first speaker identification unit 406 obtains the feature amounts of two pieces of male voice data from the feature amounts of the plurality of pieces of male voice data extracted by the feature amount extraction unit 403 .

[0309] Next, in step S159, first speaker identification unit 406 calculates the similarity between the two male voice data by inputting the feature values ​​of the acquired two male voice data into the first speaker identification model. The two male voice data may be either two male voice data spoken by a single speaker or two male voice data spoken by two speakers. In this case, the similarity between the two male voice data spoken by a single speaker is higher than the similarity between the two male voice data spoken by two speakers.

[0310] Next, in step S160, first speaker identification unit 406 determines whether similarities have been calculated for all combinations of male voice data. If it is determined that similarities have not been calculated for all combinations of male voice data (No in step S160), processing returns to step S158. First speaker identification unit 406 then obtains, from feature extraction unit 403, feature quantities for the two male voice data sets for which similarities have not been calculated, from the feature quantities of the plurality of male voice data sets.

[0311] On the other hand, when it is determined that the similarities of all combinations of male voice data have been calculated (yes in step S160), in step S161, the first threshold calculation unit 407 calculates a first threshold value that can identify the similarities of two male voice data of the same speaker and the similarities of two male voice data of different speakers by performing regression analysis on the multiple similarities calculated by the first speaker recognition unit 406.

[0312] Next, in step S162 , the first threshold calculation unit 407 stores the calculated first threshold in the threshold storage unit 408 .

[0313] The processing of steps S163 to S168 is Figure 9 and Figure 10 The processes of steps S67 to S72 shown are the same, so their description is omitted.

[0314] Next, in step S169 , the second speaker recognition unit 416 acquires the second speaker recognition model from the second speaker recognition model storage unit 415 .

[0315] Next, in step S170 , the second speaker recognition unit 416 obtains the feature values ​​of two pieces of female voice data from the feature values ​​of the plurality of pieces of female voice data extracted by the feature value extraction unit 413 .

[0316] Next, in step S171, the second speaker recognition unit 416 calculates the similarity between the two female voice data by inputting the feature values ​​of the two acquired female voice data into the second speaker recognition model. The two female voice data may be either two female voice data spoken by a single speaker or two female voice data spoken by two speakers. In this case, the similarity between the two female voice data spoken by a single speaker is higher than the similarity between the two female voice data spoken by two speakers.

[0317] Next, in step S172, second speaker recognition unit 416 determines whether similarities have been calculated for all combinations of female voice data. If it is determined that similarities have not been calculated for all combinations of female voice data (No in step S172), processing returns to step S170. Second speaker recognition unit 416 then obtains, from feature extraction unit 413, feature quantities for two female voice data sets for which similarities have not been calculated, from the feature quantities of the plurality of female voice data sets.

[0318] On the other hand, when it is determined that the similarities of all combinations of female voice data have been calculated (yes in step S172), in step S173, the second threshold calculation unit 417 calculates the second threshold value that can identify the similarities of two female voice data of the same speaker and the similarities of two female voice data of different speakers by performing regression analysis on the multiple similarities calculated by the second speaker recognition unit 416.

[0319] Next, in step S174 , the second threshold calculation unit 417 stores the calculated second threshold in the threshold storage unit 408 .

[0320] Next, the operation of the speaker identification process performed by the speaker identification device 22 according to the third embodiment will be described.

[0321] Figure 23 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the third embodiment. Figure 24 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the third embodiment.

[0322] The processing of steps S181 to S191 is Figure 16 and Figure 17 The processes of steps S131 to S141 shown are the same, so their description is omitted.

[0323] Next, in step S192, the similarity correction unit 233 obtains the first threshold or the second threshold from the threshold storage unit 223. If the model selection unit 217 has selected the first speaker recognition model, the similarity correction unit 233 obtains the first threshold from the threshold storage unit 223. If the model selection unit 217 has selected the second speaker recognition model, the similarity correction unit 233 obtains the second threshold from the threshold storage unit 223.

[0324] Next, in step S193, the similarity correction unit 233 uses the acquired first threshold or second threshold to correct the similarity calculated by the similarity calculation unit 2311. At this time, the similarity correction unit 233 subtracts the first threshold or second threshold from the similarity calculated by the similarity calculation unit 2311.

[0325] Next, in step S194, the similarity determination unit 2322 determines whether the similarity corrected by the similarity correction unit 233 is greater than a third threshold value. The third threshold value is, for example, 0. If the corrected similarity is greater than 0, the similarity determination unit 2322 determines that the speech data to be recognized is consistent with the pre-registered registration speech data. If the corrected similarity is less than 0, the similarity determination unit 2322 determines that the speech data to be recognized is inconsistent with the pre-registered registration speech data.

[0326] Here, when the similarity correcting unit 233 determines that the similarity corrected is greater than the third threshold (Yes in step S194 ), the similarity determining unit 2322 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data in step S195 .

[0327] On the other hand, when the similarity correcting unit 233 determines that the similarity corrected is below the third threshold (No in step S194), the similarity determining unit 2322 determines in step S196 that the speaker of the recognition target speech data is not the speaker of the registered speech data.

[0328] The processing of step S197 is the same as Figure 17 The processing of step S145 shown is the same, so the description is omitted.

[0329] When using two different first and second speaker recognition models, the output value ranges of the first and second speaker recognition models may differ. To address this issue, in this third embodiment, during registration, a first threshold and a second threshold are calculated for each of the first and second speaker recognition models, respectively, to identify the same speaker. Furthermore, during speaker recognition, the similarity between the calculated speech data to be recognized and the registered speech data is corrected by subtracting the first or second threshold from the calculated similarity. Furthermore, by comparing the corrected similarity with a third threshold shared by both the first and second speaker recognition models, the speaker of the speech data to be recognized can be identified with higher accuracy.

[0330] (Implementation Method 4)

[0331] The longer a speaker speaks, the more information is conveyed, making it easier to identify the speaker. The similarity between the registered voice data and the target voice data for the same person tends to be higher. On the other hand, the shorter a speaker speaks, the less information is conveyed, making it more difficult to identify the speaker. Even for the same person, the similarity between the registered voice data and the target voice data for recognition may be lower. Therefore, if a speaker recognition model that has been machine-learned using speech data from longer speech durations identifies the speaker of target voice data from shorter speech durations, the accuracy of speaker recognition may be reduced.

[0332] In this regard, the speaker recognition method of implementation mode 4 allows a computer to execute: obtaining recognition object voice data; obtaining pre-registered registration voice data; extracting feature quantities of the recognition object voice data; extracting feature quantities of the registration voice data; when the speaking time of at least one of the speakers of the recognition object voice data and the registration voice data is longer than a specified time, selecting a third speaker recognition model that has been machine-learned using voice data with a speaking time longer than a specified time in order to recognize speakers with a speaking time longer than a specified time; when the speaking time of at least one of the speakers of the recognition object voice data and the registration voice data is shorter than a specified time, selecting a fourth speaker recognition model that has been machine-learned using voice data with a speaking time shorter than a specified time in order to recognize speakers with a speaking time shorter than a specified time; and recognizing the speaker of the recognition object voice data by inputting the feature quantities of the recognition object voice data and the feature quantities of the registration voice data into any one of the selected third speaker recognition model and the fourth speaker recognition model.

[0333] Figure 25 This is a diagram showing the configuration of a speaker recognition system according to a fourth embodiment of the present invention.

[0334] Figure 25 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 24. The speaker recognition device 24 may or may not include the microphone 1.

[0335] In addition, in this embodiment 4, the same reference numerals are attached to the same configurations as those in the embodiment 1, and description thereof will be omitted.

[0336] The speaker recognition device 24 includes a registration object voice data acquisition unit 201, a speech time measurement unit 207, a registration unit 2064, a recognition object voice data acquisition unit 211, a registration voice data storage unit 2124, a registration voice data acquisition unit 213, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 2164, a model selection unit 2174, a speaker recognition unit 2184 and a recognition result output unit 219.

[0337] The utterance time measuring unit 207 measures the utterance time of the registration target voice data acquired by the registration target voice data acquiring unit 201. The utterance time is the time from the time when the registration target voice data acquiring unit 201 starts acquiring the registration target voice data to the time when the acquisition of the registration target voice data is completed.

[0338] The registration unit 2064 registers, as registration voice data, voice data to be registered that corresponds to the utterance time information indicating the utterance time measured by the utterance time measurement unit 207. The registration unit 2064 registers the registration voice data in the registration voice data storage unit 2124.

[0339] The speaker recognition device 24 may further include an input receiving unit for receiving information about the speaker of the voice data to be registered. Furthermore, the registration unit 2064 may register the registered voice data and the speaker information in the registered voice data storage unit 2124 in association with each other. The speaker information may include, for example, the speaker's name.

[0340] The registered voice data storage unit 2124 stores the registered voice data corresponding to the utterance time information. The registered voice data storage unit 2124 stores a plurality of registered voice data.

[0341] The speaker recognition model storage unit 2164 pre-stores: a third speaker recognition model, which is machine-learned using speech data with a speaking time of at least a predetermined time, to identify speakers who speak for at least a predetermined time; and a fourth speaker recognition model, which is machine-learned using speech data with a speaking time of less than a predetermined time, to identify speakers who speak for less than a predetermined time. The speaker recognition model storage unit 2164 pre-stores the third and fourth speaker recognition models generated by the speaker recognition model generation device 44, described later. The methods for generating the third and fourth speaker recognition models will be described later.

[0342] If at least one of the speaker of the target speech data and the speaker of the registered speech data spoke for a predetermined time or longer, the model selection unit 2174 selects a third speaker recognition model that has been machine-learned using speech data with a predetermined time or longer in order to recognize speakers with a predetermined time or longer. Furthermore, if at least one of the speaker of the target speech data and the speaker of the registered speech data spoke for a predetermined time or shorter, the model selection unit 2174 selects a fourth speaker recognition model that has been machine-learned using speech data with a predetermined time or shorter in order to recognize speakers with a predetermined time or shorter.

[0343] In this fourth embodiment, the model selection unit 2174 selects the third speaker recognition model when the speaker's speaking time in the registered voice data exceeds a predetermined time, and selects the fourth speaker recognition model when the speaker's speaking time in the registered voice data is shorter than the predetermined time. The registered voice data is previously associated with the speaking time. Therefore, the model selection unit 2174 selects the third speaker recognition model when the speaking time corresponding to the registered voice data exceeds the predetermined time, and selects the fourth speaker recognition model when the speaking time corresponding to the registered voice data is shorter than the predetermined time. The predetermined time is, for example, 60 seconds.

[0344] The speaker recognition unit 2184 recognizes the speaker of the recognition target speech data by inputting the feature value of the recognition target speech data and the feature value of the registration speech data into either the third speaker recognition model or the fourth speaker recognition model selected by the model selection unit 2174 .

[0345] The speaker recognition unit 2184 includes a similarity calculation unit 2314 and a similarity determination unit 232 .

[0346] The similarity calculation unit 2314 obtains the similarity between the recognition target speech data and the multiple registration speech data from any one of the third speaker recognition model and the fourth speaker recognition model by inputting the feature value of the recognition target speech data and the respective feature values ​​of the multiple registration speech data into any one of the selected third speaker recognition model and the fourth speaker recognition model.

[0347] Next, a speaker recognition model generation device according to a fourth embodiment of the present invention will be described.

[0348] Figure 26 This is a diagram showing the configuration of a speaker recognition model generation device according to a fourth embodiment of the present invention.

[0349] Figure 26 The speaker recognition model generation device 44 shown includes a long-time speech data storage unit 421, a long-time speech data acquisition unit 422, a feature extraction unit 423, a third speaker recognition model generation unit 424, a third speaker recognition model storage unit 425, a short-time speech data storage unit 431, a short-time speech data acquisition unit 432, a feature extraction unit 433, a fourth speaker recognition model generation unit 434 and a fourth speaker recognition model storage unit 435.

[0350] The long-duration speech data acquisition unit 422, feature extraction unit 423, third speaker recognition model generation unit 424, short-duration speech data acquisition unit 432, feature extraction unit 433, and fourth speaker recognition model generation unit 434 are implemented by a processor. The long-duration speech data storage unit 421, third speaker recognition model storage unit 425, short-duration speech data storage unit 431, and fourth speaker recognition model storage unit 435 are implemented by a memory.

[0351] The long-duration speech data storage unit 421 stores a plurality of long-duration speech data items, each of which is assigned a speaker identification tag for speaker identification and has a utterance duration of at least a predetermined time. Long-duration speech data items are speech data that has a utterance duration of at least a predetermined time. The long-duration speech data storage unit 421 stores a plurality of different long-duration speech data items for each of the plurality of speakers.

[0352] The long-duration speech data acquisition unit 422 acquires a plurality of long-duration speech data items, each of which is assigned a speaker identification tag for identifying a speaker, from the long-duration speech data storage unit 421. In Embodiment 4, the long-duration speech data acquisition unit 422 acquires a plurality of long-duration speech data items from the long-duration speech data storage unit 421. However, the present invention is not particularly limited to this embodiment. Alternatively, the plurality of long-duration speech data items may be acquired (received) from an external device via a network.

[0353] The feature extraction unit 423 extracts feature values ​​from the plurality of long-duration speech data acquired by the long-duration speech data acquisition unit 422. The feature values ​​are, for example, i-vectors.

[0354] The third speaker recognition model generation unit 424 uses the feature values ​​of the first and second long speech data sets, as well as the similarity between the speaker recognition labels of the first and second long speech data sets, as training data, and generates a third speaker recognition model through machine learning, which takes the feature values ​​of the two speech data sets as input and outputs the similarity between the two speech data sets. For example, the third speaker recognition model performs machine learning so that if the speaker recognition labels of the first and second long speech data sets are the same, the third speaker recognition model outputs the highest similarity; if the speaker recognition labels of the first and second long speech data sets are different, the third speaker recognition model outputs the lowest similarity.

[0355] The third speaker recognition model uses a model using PLDA. The PLDA model automatically selects features effective for speaker recognition from a 400-dimensional i-vector (feature quantity) and calculates the log-likelihood ratio as the similarity.

[0356] Other examples of machine learning include supervised learning, which uses teacher data that labels input information (output information) to learn the relationship between input and output; unsupervised learning, which constructs data structures based solely on unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which uses trial and error to learn behaviors that maximize rewards. Specific machine learning methods include neural networks (including deep learning using multilayer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). Any of these specific examples can be used in the machine learning of the third speaker recognition model.

[0357] The third speaker recognition model storage unit 425 stores the third speaker recognition model generated by the third speaker recognition model generation unit 424 .

[0358] The short-term speech data storage unit 431 stores a plurality of short-term speech data items, each of which is assigned a speaker identification tag for speaker identification and has a utterance duration shorter than a predetermined time. Short-term speech data items are speech data whose utterance duration is shorter than the predetermined time. The short-term speech data storage unit 431 stores a plurality of different short-term speech data items for each of the plurality of speakers.

[0359] The short-time speech data acquisition unit 432 acquires a plurality of short-time speech data items, each of which has been assigned a speaker identification tag for identifying a speaker, from the short-time speech data storage unit 431. In Embodiment 4, the short-time speech data acquisition unit 432 acquires a plurality of short-time speech data items from the short-time speech data storage unit 431. However, the present invention is not particularly limited to this embodiment. Alternatively, the plurality of short-time speech data items may be acquired (received) from an external device via a network.

[0360] The feature amount extraction unit 433 extracts feature amounts from the plurality of short-time speech data acquired by the short-time speech data acquisition unit 432. The feature amount is, for example, an i-vector.

[0361] The fourth speaker recognition model generation unit 434 uses the feature values ​​of the first and second short-term speech data, as well as the similarity between the speaker recognition labels of the first and second short-term speech data, as training data, and generates a fourth speaker recognition model through machine learning, which takes the feature values ​​of the two speech data as input and outputs the similarity between the two speech data. For example, the fourth speaker recognition model performs machine learning so that if the speaker recognition labels of the first and second short-term speech data are the same, the highest similarity is output; if the speaker recognition labels of the first and second short-term speech data are different, the lowest similarity is output.

[0362] The fourth speaker recognition model uses a model using PLDA. The PLDA model automatically selects features effective for speaker recognition from a 400-dimensional i-vector (feature quantity) and calculates the log-likelihood ratio as the similarity.

[0363] Other examples of machine learning include supervised learning, which uses teacher data that labels input information (output information) to learn the relationship between input and output; unsupervised learning, which constructs data structures based solely on unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which uses trial and error to learn behaviors that maximize rewards. Specific machine learning methods include neural networks (including deep learning using multilayer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). Any of these specific examples can be used in the machine learning of the fourth speaker recognition model.

[0364] The fourth speaker recognition model storage unit 435 stores the fourth speaker recognition model generated by the fourth speaker recognition model generation unit 434 .

[0365] Furthermore, the speaker recognition model generation device 44 may transmit the third speaker recognition model stored in the third speaker recognition model storage unit 425 and the fourth speaker recognition model stored in the fourth speaker recognition model storage unit 435 to the speaker recognition device 24. The speaker recognition device 24 may store the received third and fourth speaker recognition models in the speaker recognition model storage unit 2164. Furthermore, when the speaker recognition device 24 is manufactured, the third and fourth speaker recognition models generated by the speaker recognition model generation device 44 may be stored in the speaker recognition device 24.

[0366] Next, the operations of the registration process and the speaker recognition process of the speaker recognition device 24 according to the fourth embodiment will be described.

[0367] Figure 27 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the fourth embodiment.

[0368] First, in step S201, the registration target voice data acquisition unit 201 acquires the registration target voice data output from the microphone 1. A speaker who wishes to register their spoken voice data speaks a predetermined passage into the microphone 1. The predetermined passage is either a passage that lasts longer than a predetermined time or a passage that lasts shorter than a predetermined time. The speaker recognition device 24 may also present the registration target speaker with multiple predetermined passages. In this case, the registration target speaker speaks the multiple presented passages.

[0369] Next, in step S202 , the utterance time measuring unit 207 measures the utterance time of the registration target voice data acquired by the registration target voice data acquiring unit 201 .

[0370] Next, in step S203 , the registration unit 2064 stores the registration target voice data associated with the utterance time information indicating the utterance time measured by the utterance time measurement unit 207 in the registration voice data storage unit 2124 as registration voice data.

[0371] Figure 28 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the fourth embodiment. Figure 29 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the fourth embodiment.

[0372] The processing of steps S211 to S214 is Figure 6 The processes of steps S31 to S34 shown are the same, so their description is omitted.

[0373] Next, in step S215 , the model selection unit 2174 acquires the utterance time corresponding to the registration voice data acquired by the registration voice data acquisition unit 213 .

[0374] Next, in step S216, the model selection unit 2174 determines whether the acquired utterance time is longer than a predetermined time. If the acquired utterance time is determined to be longer than the predetermined time (yes in step S216), the model selection unit 2174 selects the third speaker identification model in step S217. The model selection unit 2174 obtains the selected third speaker identification model from the speaker identification model storage unit 2164 and outputs the obtained third speaker identification model to the similarity calculation unit 2314.

[0375] On the other hand, if the acquired utterance time is determined to be less than the predetermined time, that is, if the acquired utterance time is determined to be less than the predetermined time (No in step S216), in step S218, the model selection unit 2174 selects the fourth speaker identification model. The model selection unit 2174 obtains the selected fourth speaker identification model from the speaker identification model storage unit 2164 and outputs the obtained fourth speaker identification model to the similarity calculation unit 2314.

[0376] Next, in step S219, the similarity calculation unit 2314 calculates the similarity between the recognition target speech data and the registration speech data by inputting the feature value of the recognition target speech data and the feature value of the registration speech data into any one of the selected third speaker recognition model and fourth speaker recognition model.

[0377] Next, in step S220, the similarity calculation unit 2314 determines whether similarities have been calculated between the target speech data and all registered speech data stored in the registered speech data storage unit 2124. If it is determined that similarities have not been calculated between the target speech data and all registered speech data (No in step S220), the process returns to step S213. The registered speech data acquisition unit 213 then acquires the registered speech data for which similarities have not been calculated from the plurality of registered speech data stored in the registered speech data storage unit 2124.

[0378] On the other hand, if it is determined that the similarities between the recognition target voice data and all the registered voice data have been calculated (Yes in step S220 ), the similarity determination unit 232 determines whether the highest similarity is greater than a threshold value in step S221 .

[0379] In addition, the processing of steps S221 to S224 is the same as Figure 7 The processes of steps S41 to S44 shown are the same, so their description is omitted.

[0380] As described above, when at least one of the speaker of the target speech data and the speaker of the registered speech data speaks for a predetermined time or longer, the speaker of the target speech data is identified by inputting the feature value of the target speech data and the feature value of the registered speech data into a third speaker recognition model that has been machine-learned using speech data that spoke for a predetermined time or longer. Furthermore, when at least one of the speaker of the target speech data and the speaker of the registered speech data speaks for a shorter time than the predetermined time, the speaker of the target speech data is identified by inputting the feature value of the target speech data and the feature value of the registered speech data into a fourth speaker recognition model that has been machine-learned using speech data that spoke for a shorter time than the predetermined time.

[0381] Therefore, the speaker of the recognition object voice data is identified using the third speaker recognition model and the fourth speaker recognition model corresponding to the length of speaking time of at least one of the recognition object voice data and the registered voice data, thereby improving the accuracy of identifying whether the speaker of the recognition object is a pre-registered speaker.

[0382] Furthermore, in this fourth embodiment, the model selection unit 2174 selects either the third speaker recognition model or the fourth speaker recognition model based on the utterance duration corresponding to the registration voice data. However, the present invention is not particularly limited to this. The speaker recognition device 24 may also include a utterance duration measurement unit that measures the utterance duration of the speaker of the registration voice data acquired by the registration voice data acquisition unit 213. The utterance duration measurement unit may also output the measured utterance duration to the model selection unit 2174. Furthermore, when measuring the utterance duration of the speaker of the registration voice data acquired by the registration voice data acquisition unit 213, the utterance duration measurement unit 207 is unnecessary, and the registration unit 2064 may simply store the registration target voice data acquired by the registration target voice data acquisition unit 201 as the registration voice data in the registration voice data storage unit 2124.

[0383] Furthermore, in this fourth embodiment, the model selection unit 2174 selects either the third or fourth speaker recognition model based on the utterance duration of the speaker in the enrollment speech data. However, the present invention is not particularly limited to this. The model selection unit 2174 may also select either the third or fourth speaker recognition model based on the utterance duration of the speaker in the target speech data. In this case, the speaker recognition device 24 may also include a utterance duration measurement unit that measures the utterance duration of the speaker in the target speech data. The utterance duration measurement unit may also measure the utterance duration of the target speech data acquired by the target speech data acquisition unit 211 and output the measured utterance duration to the model selection unit 2174. The model selection unit 2174 selects the third speaker recognition model if the utterance duration of the speaker in the target speech data is longer than a predetermined time, and selects the fourth speaker recognition model if the utterance duration of the speaker in the target speech data is shorter than the predetermined time. The predetermined time is, for example, 30 seconds. Furthermore, when measuring the speaking time of the speaker of the recognition target speech data, the speaking time measuring unit 207 is not required, and the registration unit 2064 may simply store the registration target speech data acquired by the registration target speech data acquiring unit 201 as registration speech data in the registration speech data storage unit 2124 .

[0384] Furthermore, in this fourth embodiment, the model selection unit 2174 may select either the third or fourth speaker recognition model based on both the utterance time of the speaker in the registration voice data and the utterance time of the speaker in the recognition target voice data. The model selection unit 2174 may select the third speaker recognition model when both the utterance time of the speaker in the registration voice data and the utterance time of the speaker in the recognition target voice data are at least a predetermined time. Furthermore, the model selection unit 2174 may select the fourth speaker recognition model when at least one of the utterance time of the speaker in the registration voice data and the utterance time of the speaker in the recognition target voice data is less than a predetermined time. The predetermined time is, for example, 20 seconds. In this case, the speaker recognition device 24 may further include a utterance time measurement unit that measures the utterance time of the speaker in the recognition target voice data. Alternatively, the speaker recognition device 24 may include, instead of the utterance time measurement unit 207, a utterance time measurement unit that measures the utterance time of the speaker in the registration voice data acquired by the registration voice data acquisition unit 213.

[0385] Furthermore, in this fourth embodiment, the model selection unit 2174 may select the third speaker recognition model when at least one of the utterance time of the speaker in the registered speech data and the utterance time of the speaker in the target speech data is longer than a predetermined time. Furthermore, the model selection unit 2174 may select the fourth speaker recognition model when both the utterance time of the speaker in the registered speech data and the utterance time of the speaker in the target speech data are shorter than a predetermined time. The predetermined time may be, for example, 100 seconds. In this case, the speaker recognition device 24 may further include a utterance time measurement unit that measures the utterance time of the speaker in the target speech data. Furthermore, the speaker recognition device 24 may include, instead of the utterance time measurement unit 207, a utterance time measurement unit that measures the utterance time of the speaker in the registered speech data acquired by the registered speech data acquisition unit 213.

[0386] Next, the operation of speaker recognition model generation processing by the speaker recognition model generation device 44 according to the fourth embodiment will be described.

[0387] Figure 30 This is a first flowchart for explaining the operation of speaker recognition model generation processing by the speaker recognition model generation device according to the fourth embodiment. Figure 31 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the fourth embodiment.

[0388] First, in step S231 , the long-time speech data acquisition unit 422 acquires, from the long-time speech data storage unit 421 , a plurality of long-time speech data items to which a speaker identification label for identifying a speaker is assigned and whose utterance time is longer than a predetermined time.

[0389] Next, in step S232 , the feature amount extraction unit 423 extracts feature amounts of the plurality of long-time speech data acquired by the long-time speech data acquisition unit 422 .

[0390] Then, in step S233, the third speaker recognition model generation unit 424 obtains the feature quantities of the first long speech data and the second long speech data among multiple long speech data and the similarity of the speaker recognition labels of the first long speech data and the second long speech data as teacher data.

[0391] Next, in step S234 , the third speaker recognition model generation unit 424 uses the acquired teacher data to machine-learn a third speaker recognition model that takes the feature quantities of the two speech data as input and outputs the similarity between the two speech data.

[0392] Next, in step S235, the third speaker recognition model generator 424 determines whether the third speaker recognition model has been trained using a combination of all the long-term speech data from the plurality of long-term speech data. If it is determined that the third speaker recognition model has not been trained using a combination of all the long-term speech data (No in step S235), the process returns to step S233. The third speaker recognition model generator 424 then obtains, as training data, the feature values ​​of the first and second long-term speech data from the plurality of long-term speech data that were not used for machine learning, as well as the similarity between the speaker recognition labels of the first and second long-term speech data.

[0393] On the other hand, when it is determined that the third speaker recognition model has been learned using the combined machine of all long-time speech data (yes in step S235), in step S236, the third speaker recognition model generation unit 424 stores the third speaker recognition model generated by machine learning in the third speaker recognition model storage unit 425.

[0394] Next, in step S237 , the short-time speech data acquisition unit 432 acquires a plurality of short-time speech data items to which a speaker identification label for identifying a speaker is assigned and whose utterance time is shorter than a predetermined time from the short-time speech data storage unit 431 .

[0395] Next, in step S238 , the feature amount extraction unit 433 extracts feature amounts of the plurality of short-time speech data acquired by the short-time speech data acquisition unit 432 .

[0396] Then, in step S239, the fourth speaker recognition model generation unit 434 obtains the feature quantities of the first short-time speech data and the second short-time speech data among multiple short-time speech data and the similarity of the speaker recognition labels of the first short-time speech data and the second short-time speech data as teacher data.

[0397] Next, in step S240 , the fourth speaker recognition model generation unit 434 uses the acquired teacher data to machine-learn a fourth speaker recognition model that takes the feature quantities of the two speech data as input and outputs the similarity between the two speech data.

[0398] Next, in step S241, the fourth speaker recognition model generation unit 434 determines whether the fourth speaker recognition model has been trained using a combination of all the short-term speech data from the plurality of short-term speech data. If it is determined that the fourth speaker recognition model has not been trained using a combination of all the short-term speech data (No in step S241), the process returns to step S239. The fourth speaker recognition model generation unit 434 then obtains, as training data, the feature values ​​of the first and second short-term speech data from the plurality of short-term speech data that were not used for machine learning, as well as the similarity between the speaker recognition labels of the first and second short-term speech data.

[0399] On the other hand, when it is determined that the fourth speaker recognition model has been machine-learned using the combination of all short-time speech data (yes in step S241), in step S242, the fourth speaker recognition model generation unit 434 stores the fourth speaker recognition model generated by machine learning in the fourth speaker recognition model storage unit 435.

[0400] (Implementation method 5)

[0401] In the fourth embodiment described above, the similarity between each of the registered voice data stored in the registered voice data storage unit 2124 and the target voice data is calculated, and the speaker of the registered voice data with the highest similarity is identified as the speaker of the target voice data. In contrast, in the fifth embodiment, the identification information of the speaker of the target voice data is input, and one piece of registered voice data that is pre-associated with the identification information is retrieved from the plurality of registered voice data stored in the registered voice data storage unit 212. The similarity between the one piece of registered voice data and the target voice data is then calculated. If the similarity exceeds a threshold, the speaker of the registered voice data is identified as the speaker of the target voice data.

[0402] Figure 32 This is a diagram showing the configuration of a speaker recognition system according to a fifth embodiment of the present invention.

[0403] Figure 32 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 25. The speaker recognition device 25 may or may not include the microphone 1.

[0404] In addition, in this embodiment 5, the same reference numerals are attached to the same structures as those in the embodiments 1 to 4, and the description thereof will be omitted.

[0405] The speaker recognition device 25 includes a registration object voice data acquisition unit 201, a speech time measurement unit 207, a registration unit 2065, a recognition object voice data acquisition unit 211, a registration voice data storage unit 2125, a registration voice data acquisition unit 2135, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 2164, a model selection unit 2174, a speaker recognition unit 2185, a recognition result output unit 219, an input acceptance unit 221 and a recognition information acquisition unit 222.

[0406] The registration unit 2065 registers, as registration voice data, voice data to be registered, which is associated with the utterance time information indicating the utterance time measured by the utterance time measurement unit 207 and the identification information acquired by the identification information acquisition unit 222. The registration unit 2065 registers the registration voice data in the registration voice data storage unit 2125.

[0407] The registered voice data storage unit 2125 stores registered voice data in association with utterance time information and identification information. The registered voice data storage unit 2125 stores a plurality of registered voice data. The plurality of registered voice data is associated with identification information for identifying the speaker of each of the plurality of registered voice data.

[0408] The registered voice data acquisition unit 2135 acquires the registered voice data corresponding to the identification information that matches the identification information acquired by the identification information acquisition unit 222 from the plurality of registered voice data registered in the registered voice data storage unit 2125 .

[0409] The speaker recognition unit 2185 includes a similarity calculation unit 2315 and a similarity determination unit 2325 .

[0410] The similarity calculation unit 2315 obtains the similarity between the recognition object speech data and the registration speech data from any one of the third speaker recognition model and the fourth speaker recognition model by inputting the feature quantity of the recognition object speech data and the feature quantity of the registration speech data into any one of the selected third speaker recognition model and the fourth speaker recognition model.

[0411] When the acquired similarity is higher than the threshold, the similarity determination unit 2325 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0412] Next, the operations of the registration process and the speaker recognition process of the speaker recognition device 25 according to the fifth embodiment will be described.

[0413] Figure 33 This is a flowchart for explaining the operation of the registration process of the speaker identification device according to the fifth embodiment of the present invention.

[0414] The processing of step S251 and step S252 is the same as Figure 15 The processing of step S121 and step S122 shown in FIG. 1 is the same, so the description is omitted. In addition, the processing of step S253 is the same as that of step S254. Figure 27 The processing of step S202 shown is the same, so the description is omitted.

[0415] Next, in step S254, the registration unit 2065 stores the registration target speech data, which corresponds to the utterance time information measured by the utterance time measurement unit 207 and the identification information acquired by the identification information acquisition unit 222, as registration speech data in the registration speech data storage unit 2125. Consequently, the registration speech data storage unit 2125 stores the registration speech data corresponding to the utterance time information and the identification information.

[0416] Figure 34 1 is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the fifth embodiment. Figure 35 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the fifth embodiment.

[0417] The processing of steps S261 to S264 is Figure 16 The processes of steps S131 to S134 shown are the same, so their description is omitted.

[0418] Next, in step S265 , the registered voice data acquisition unit 2135 acquires the registered voice data corresponding to the identification information that matches the identification information acquired by the identification information acquisition unit 222 from the plurality of registered voice data registered in the registered voice data storage unit 2125 .

[0419] The processing of steps S266 to S270 is the same as Figure 28 The processes of steps S214 to S218 shown are the same, so their description is omitted.

[0420] Then, in step S271, the similarity calculation unit 2315 obtains the similarity between the recognition object speech data and the registration speech data from any one of the third speaker recognition model and the fourth speaker recognition model by inputting the feature value of the recognition object speech data and the feature value of the registration speech data into any one of the selected third speaker recognition model and the fourth speaker recognition model.

[0421] Next, in step S272 , the similarity determination unit 2325 determines whether the similarity calculated by the similarity calculation unit 2315 is greater than a threshold value.

[0422] Here, when it is determined that the similarity calculated by the similarity calculation unit 2315 is greater than the threshold (Yes in step S272 ), in step S273 , the similarity determination unit 2325 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0423] On the other hand, when the similarity calculation unit 2315 determines that the similarity calculated is below the threshold (No in step S272), the similarity determination unit 2325 determines in step S274 that the speaker of the recognition target speech data is not the speaker of the registered speech data.

[0424] The processing of step S275 is the same as Figure 7 The processing of step S44 shown is the same, so the description is omitted.

[0425] As described above, in this fifth embodiment, only the similarity between the registration voice data corresponding to the recognition information and the recognition target voice data is calculated. Therefore, compared to the fourth embodiment, which calculates multiple similarities between each of multiple registration voice data and the recognition target voice data, the fifth embodiment can reduce the processing load for calculating similarities.

[0426] (Implementation Method 6)

[0427] In the aforementioned embodiments 4 and 5, either the third or fourth speaker recognition model is selected based on the speaker's utterance time in the enrollment speech data. However, when using two different third and fourth speaker recognition models, the output value ranges of the third and fourth speaker recognition models may differ. To address this issue, in embodiment 6, when enrolling speech data, a third threshold and a fourth threshold for identifying the same speaker are calculated for each of the third and fourth speaker recognition models, and the calculated third and fourth thresholds are stored. Furthermore, when recognizing speech data, the similarity between the target speech data and the enrollment speech data is corrected by subtracting the third or fourth threshold from the calculated similarity. The corrected similarity is then compared with the fifth threshold, which is common to both the third and fourth speaker recognition models, to identify the speaker of the target speech data.

[0428] First, a speaker recognition model generation device according to a sixth embodiment of the present invention will be described.

[0429] Figure 36 This is a diagram showing the configuration of a speaker recognition model generation device according to a sixth embodiment of the present invention.

[0430] Figure 36The speaker recognition model generation device 46 shown includes a long-time speech data storage unit 421, a long-time speech data acquisition unit 422, a feature extraction unit 423, a third speaker recognition model generation unit 424, a third speaker recognition model storage unit 425, a third speaker recognition unit 426, a third threshold calculation unit 427, a threshold storage unit 428, a short-time speech data storage unit 431, a short-time speech data acquisition unit 432, a feature extraction unit 433, a fourth speaker recognition model generation unit 434, a fourth speaker recognition model storage unit 435, a fourth speaker recognition unit 436 and a fourth threshold calculation unit 437.

[0431] In addition, in this embodiment 6, the same reference numerals are attached to the same structures as those in the embodiments 1 to 5, and the description thereof will be omitted.

[0432] The third speaker recognition unit 426 inputs all combinations of feature quantities of two speech data among a plurality of speech data with a utterance time of at least a predetermined time into the third speaker recognition model, and obtains similarities of each of the plurality of combinations of the two speech data from the third speaker recognition model.

[0433] The third threshold calculation unit 427 calculates a third threshold that can distinguish the similarity between two speech data items of the same speaker and the similarity between two speech data items of different speakers. The third threshold calculation unit 427 calculates the third threshold by performing regression analysis on the multiple similarities calculated by the third speaker identification unit 426.

[0434] The fourth speaker recognition unit 436 inputs all combinations of feature quantities of two speech data among a plurality of speech data whose utterance time is shorter than a predetermined time into the fourth speaker recognition model, and obtains similarities of each of the plurality of combinations of the two speech data from the fourth speaker recognition model.

[0435] The fourth threshold calculation unit 437 calculates a fourth threshold that enables the distinction between the similarity between two speech data items of the same speaker and the similarity between two speech data items of different speakers. The fourth threshold calculation unit 437 calculates the fourth threshold by performing regression analysis on the multiple similarities calculated by the fourth speaker identification unit 436.

[0436] The threshold storage unit 428 stores the third threshold calculated by the third threshold calculation unit 427 and the fourth threshold calculated by the fourth threshold calculation unit 437 .

[0437] Next, a speaker recognition system according to a sixth embodiment of the present invention will be described.

[0438] Figure 37 This is a diagram showing the configuration of a speaker recognition system according to a sixth embodiment of the present invention.

[0439] Figure 37 The speaker recognition system shown includes a microphone 1 and a speaker recognition device 26. The speaker recognition device 26 may or may not include the microphone 1.

[0440] In addition, in this embodiment 6, the same reference numerals are attached to the same structures as those in the embodiments 1 to 5, and the description thereof will be omitted.

[0441] The speaker recognition device 26 includes a registration object voice data acquisition unit 201, a speech time measurement unit 207, a registration unit 2065, a recognition object voice data acquisition unit 211, a registration voice data storage unit 2125, a registration voice data acquisition unit 2135, a feature value extraction unit 214, a feature value extraction unit 215, a speaker recognition model storage unit 2164, a model selection unit 2174, a speaker recognition unit 2186, a recognition result output unit 219, an input acceptance unit 221, a recognition information acquisition unit 222 and a threshold storage unit 2236.

[0442] The speaker recognition unit 2186 includes a similarity calculation unit 2315 , a similarity correction unit 2336 , and a similarity determination unit 2326 .

[0443] When the similarity is obtained from the third speaker recognition model, the similarity correction unit 2336 subtracts the third threshold from the obtained similarity. Furthermore, when the similarity is obtained from the fourth speaker recognition model, the similarity correction unit 2336 subtracts the fourth threshold from the obtained similarity. When the similarity calculation unit 2315 calculates similarity using the third speaker recognition model, the similarity correction unit 2336 reads the third threshold from the threshold storage unit 2236 and subtracts the third threshold from the calculated similarity. Furthermore, when the similarity calculation unit 2315 calculates similarity using the fourth speaker recognition model, the similarity correction unit 2336 reads the fourth threshold from the threshold storage unit 2236 and subtracts the fourth threshold from the calculated similarity.

[0444] The threshold storage unit 2236 pre-stores a third threshold value for correcting the similarity calculated using the third speaker recognition model and a fourth threshold value for correcting the similarity calculated using the fourth speaker recognition model. The threshold storage unit 2236 pre-stores the third and fourth threshold values ​​generated by the speaker recognition model generation device 46.

[0445] Furthermore, speaker recognition model generation device 46 may transmit the third and fourth threshold values ​​stored in threshold storage unit 428 to speaker recognition device 26. Speaker recognition device 26 may store the received third and fourth threshold values ​​in threshold storage unit 2236. Furthermore, when speaker recognition device 26 is manufactured, the third and fourth threshold values ​​generated by speaker recognition model generation device 46 may be stored in threshold storage unit 2236.

[0446] When the similarity obtained by subtracting the third threshold or the fourth threshold by the similarity correction unit 2336 is higher than the fifth threshold, the similarity determination unit 2326 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data.

[0447] Next, the operation of speaker recognition model generation processing by speaker recognition model generation device 46 according to the sixth embodiment will be described.

[0448] Figure 38 This is a first flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the sixth embodiment. Figure 39 This is a second flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the sixth embodiment. Figure 40 This is a third flowchart for explaining the operation of the speaker recognition model generation process of the speaker recognition model generation device according to the sixth embodiment.

[0449] The processing of steps S281 to S286 is Figure 30 The processes of steps S231 to S236 shown are the same, so their description is omitted.

[0450] Next, in step S287 , the third speaker recognition unit 426 acquires the third speaker recognition model from the third speaker recognition model storage unit 425 .

[0451] Next, in step S288 , the third speaker recognition unit 426 obtains two feature quantities of the long-duration speech data from the feature quantities of the plurality of long-duration speech data extracted by the feature quantity extraction unit 423 .

[0452] Next, in step S289, the third speaker recognition unit 426 calculates the similarity between the two long speech data sets by inputting the feature values ​​of the two acquired long speech data sets into the third speaker recognition model. The two long speech data sets may be either two long speech data sets spoken by a single speaker or two long speech data sets spoken by two speakers. In this case, the similarity between the two long speech data sets spoken by a single speaker is higher than the similarity between the two long speech data sets spoken by two speakers.

[0453] Next, in step S290, the third speaker recognition unit 426 determines whether similarities have been calculated for all combinations of long-duration speech data. If it is determined that similarities have not been calculated for all combinations of long-duration speech data (No in step S290), the process returns to step S288. The third speaker recognition unit 426 then obtains, from the feature extraction unit 423, feature quantities for two long-duration speech data sets for which similarities have not been calculated.

[0454] On the other hand, when it is determined that the similarity of the combination of all long-time speech data has been calculated (yes in step S290), in step S291, the third threshold calculation unit 427 calculates the third threshold value that can identify the similarity of two long-time speech data of the same speaker and the similarity of two long-time speech data of different speakers by performing regression analysis on the multiple similarities calculated by the third speaker recognition unit 426.

[0455] Next, in step S292 , the third threshold calculation unit 427 stores the calculated third threshold in the threshold storage unit 428 .

[0456] The processing of steps S293 to S298 is the same as Figure 30 and Figure 31 The processing of steps S237 to S242 shown is the same, so the description is omitted.

[0457] Next, in step S299 , the fourth speaker recognition unit 436 acquires the fourth speaker recognition model from the fourth speaker recognition model storage unit 435 .

[0458] Next, in step S300 , the fourth speaker recognition unit 436 obtains the feature values ​​of two pieces of short-term speech data from the feature values ​​of the plurality of pieces of short-term speech data extracted by the feature value extraction unit 433 .

[0459] Next, in step S301, the fourth speaker recognition unit 436 calculates the similarity between the two acquired short-term speech data by inputting the feature values ​​of the two short-term speech data into the fourth speaker recognition model. The two short-term speech data may be either two short-term speech data spoken by a single speaker or two short-term speech data spoken by two speakers. In this case, the similarity between the two short-term speech data is higher when the two short-term speech data are long-term speech data spoken by a single speaker than when the two short-term speech data are short-term speech data spoken by two speakers.

[0460] Next, in step S302, the fourth speaker recognition unit 436 determines whether similarities have been calculated for all combinations of short-duration speech data. If it is determined that similarities have not been calculated for all combinations of short-duration speech data (No in step S302), processing returns to step S300. The fourth speaker recognition unit 436 then obtains, from the feature extraction unit 433, feature quantities for two short-duration speech data sets for which similarities have not been calculated, from among the feature quantities of the plurality of short-duration speech data sets.

[0461] On the other hand, when it is determined that the similarity of the combination of all short-time speech data has been calculated (yes in step S302), in step S303, the fourth threshold calculation unit 437 calculates the fourth threshold value that can identify the similarity of two short-time speech data of the same speaker and the similarity of two short-time speech data of different speakers by performing regression analysis on the multiple similarities calculated by the fourth speaker recognition unit 436.

[0462] Next, in step S304 , the fourth threshold calculation unit 437 stores the calculated fourth threshold in the threshold storage unit 428 .

[0463] Next, the operation of speaker identification processing by the speaker identification device 26 according to the sixth embodiment will be described.

[0464] Figure 41 This is a first flowchart for explaining the operation of speaker recognition processing by the speaker recognition device according to the sixth embodiment. Figure 42 This is a second flowchart for explaining the operation of the speaker recognition process of the speaker recognition device according to the sixth embodiment.

[0465] The processing of steps S311 to S321 is Figure 34 and Figure 35 The processes of steps S261 to S271 shown are the same, so their description is omitted.

[0466] Next, in step S322, the similarity correction unit 2336 obtains the third threshold or the fourth threshold from the threshold storage unit 2236. In this case, if the model selection unit 2174 selected the third speaker recognition model, the similarity correction unit 2336 obtains the third threshold from the threshold storage unit 2236. Furthermore, if the model selection unit 2174 selected the fourth speaker recognition model, the similarity correction unit 2336 obtains the fourth threshold from the threshold storage unit 2236.

[0467] Next, in step S323, the similarity correction unit 2336 uses the acquired third threshold or fourth threshold to correct the similarity calculated by the similarity calculation unit 2315. At this time, the similarity correction unit 2336 subtracts the third threshold or fourth threshold from the similarity calculated by the similarity calculation unit 2315.

[0468] Next, in step S324, the similarity determination unit 2326 determines whether the similarity corrected by the similarity correction unit 2336 is greater than a fifth threshold value. The fifth threshold value is, for example, 0. If the corrected similarity is greater than 0, the similarity determination unit 2326 determines that the speech data to be recognized is consistent with the pre-registered registration speech data. If the corrected similarity is less than 0, the similarity determination unit 2326 determines that the speech data to be recognized is inconsistent with the pre-registered registration speech data.

[0469] Here, when the similarity correcting unit 2336 determines that the similarity corrected is greater than the fifth threshold (Yes in step S324 ), the similarity determining unit 2326 recognizes the speaker of the registration voice data as the speaker of the recognition target voice data in step S325 .

[0470] On the other hand, when the similarity correcting unit 2336 determines that the similarity corrected is below the fifth threshold (No in step S324), the similarity determining unit 2326 determines in step S326 that the speaker of the recognition target speech data is not the speaker of the registered speech data.

[0471] The processing of step S327 is the same as Figure 35 The processing of step S275 shown is the same, so the description is omitted.

[0472] When using two different third and fourth speaker recognition models, the output value ranges of the third and fourth speaker recognition models may differ. To address this issue, in this sixth embodiment, during registration, a third threshold and a fourth threshold that enable identification of the same speaker are calculated for each of the third and fourth speaker recognition models. Furthermore, during speaker recognition, the similarity between the calculated speech data to be recognized and the registered speech data is corrected by subtracting the third or fourth threshold from the calculated similarity. Furthermore, by comparing the corrected similarity with the fifth threshold shared by both the third and fourth speaker recognition models, the speaker of the speech data to be recognized can be identified with higher accuracy.

[0473] In addition, in each of the above embodiments, each component may be formed by dedicated hardware or implemented by executing a software program suitable for each component. Each component may be implemented by a program execution unit such as a CPU or a processor reading and executing a software program stored in a storage medium such as a hard disk or semiconductor memory.

[0474] Part or all of the functions of the devices described in the embodiments of the present invention are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These functions can be integrated individually into a single chip, or a single chip can include some or all of the functions. Furthermore, integrated circuitry is not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. Alternatively, a programmable FPGA (Field Programmable Gate Array) or a reconfigurable processor, which allows the connections and settings of circuit elements within the LSI to be reconfigured after LSI manufacturing, can be utilized.

[0475] Furthermore, part or all of the functions of the apparatus according to the embodiment of the present invention may be realized by executing a program on a processor such as a CPU.

[0476] In addition, all the numbers used above are exemplified to specifically describe the present invention, and the present invention is not limited to the exemplified numbers.

[0477] In addition, the order in which the steps shown in the above flowcharts are executed is an example order for the purpose of specifically explaining the present invention, and an order other than the above order can be adopted to the extent that the same effect is achieved. In addition, some of the above steps can also be executed simultaneously (in parallel) with other steps.

[0478] Industrial applicability

[0479] The technology of the present invention can improve the accuracy of identifying whether a speaker of an identification target is a pre-registered speaker, and therefore has practical value as a technology for identifying speakers.

Claims

1. A speaker recognition method, characterized in that: Have the computer perform the following steps: Obtaining the voice data of the recognition object; Acquire pre-registered registration voice data; Extracting feature quantities of the speech data of the recognition object; extracting a feature value of the enrollment voice data; If the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is male, a first speaker recognition model that has been machine-learned using male voice data for recognizing male speakers is selected; and if the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is female, a second speaker recognition model that has been machine-learned using female voice data for recognizing female speakers is selected; and The speaker of the recognition target speech data is recognized by inputting the feature quantity of the recognition target speech data and the feature quantity of the enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, When identifying the speaker, By inputting the feature quantity of the recognition target speech data and each feature quantity of the plurality of enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, a similarity between the recognition target speech data and each of the plurality of enrollment speech data is obtained from any one of the first speaker recognition model and the second speaker recognition model, Identify the speaker of the registered voice data with the highest acquired similarity as the speaker of the recognition target voice data, During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. When identifying the speaker, when the similarity is obtained from the first speaker recognition model, the first threshold is subtracted from the obtained similarity; when the similarity is obtained from the second speaker recognition model, the second threshold is subtracted from the obtained similarity.

2. A speaker recognition method, characterized in that: Have the computer perform the following steps: Obtaining the voice data of the recognition object; Acquire pre-registered registration voice data; Extracting feature quantities of the speech data of the recognition object; extracting a feature value of the enrollment voice data; If the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is male, a first speaker recognition model that has been machine-learned using male voice data for recognizing male speakers is selected; and if the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is female, a second speaker recognition model that has been machine-learned using female voice data for recognizing female speakers is selected; and The speaker of the recognition target speech data is recognized by inputting the feature quantity of the recognition target speech data and the feature quantity of the enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, The plurality of registered voice data are associated with identification information for identifying a speaker of each of the plurality of registered voice data. further acquiring identification information of a speaker for identifying the recognition target voice data; When acquiring the registered voice data, acquiring the registered voice data corresponding to the identification information that is consistent with the acquired identification information from the plurality of registered voice data, When identifying the speaker, By inputting the feature quantity of the recognition target speech data and the feature quantity of the enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, a similarity between the recognition target speech data and the enrollment speech data is obtained from any one of the first speaker recognition model and the second speaker recognition model, If the obtained similarity is higher than a threshold, the speaker of the registration voice data is identified as the speaker of the recognition target voice data. During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. When identifying the speaker, when the similarity is obtained from the first speaker recognition model, the first threshold is subtracted from the obtained similarity; when the similarity is obtained from the second speaker recognition model, the second threshold is subtracted from the obtained similarity.

3. The speaker recognition method according to claim 1 or 2, characterized in that: When selecting the first speaker recognition model or the second speaker recognition model, When the gender of the speaker of the registration voice data is male, selecting the first speaker recognition model, When the gender of the speaker of the registration voice data is female, the second speaker recognition model is selected.

4. The speaker recognition method according to claim 1 or 2, characterized in that: Also obtain the registration object voice data; further extracting a feature quantity of the voice data of the enrollment object; further identifying the gender of a speaker of the registration object voice data by using the feature amount of the registration object voice data; The registration target voice data corresponding to the identified gender is also registered as the registration voice data.

5. The speaker recognition method according to claim 4, characterized in that In identifying said gender, Obtain a gender recognition model that uses machine learning to identify the speaker's gender using male and female voice data. The gender of the speaker of the registration object voice data is identified by inputting the feature amount of the registration object voice data into the gender identification model.

6. The speaker recognition method according to claim 5, characterized in that In identifying said gender, By inputting the feature amount of the voice data of the registration object and each feature amount of a plurality of pre-stored voice data of males into the gender recognition model, the similarity between the voice data of the registration object and each of the plurality of voice data of males is obtained from the gender recognition model, The average of the obtained multiple similarities is used as the average male similarity to calculate, By inputting the feature amount of the voice data of the registration object and each feature amount of a plurality of pre-stored voice data of a female into the gender recognition model, the similarity between the voice data of the registration object and each of the plurality of voice data of the female is obtained from the gender recognition model, The average of the obtained multiple similarities is used as the average female similarity to calculate, If the average male similarity is higher than the average female similarity, the gender of the speaker of the registration object voice data is identified as male, If the average male similarity is lower than the average female similarity, the gender of the speaker of the registration target voice data is identified as female.

7. The speaker recognition method according to claim 5, characterized in that In identifying said gender, By inputting the feature amount of the voice data of the registration object and each feature amount of a plurality of pre-stored voice data of males into the gender recognition model, the similarity between the voice data of the registration object and each of the plurality of voice data of males is obtained from the gender recognition model, The maximum value among the obtained multiple similarities is calculated as the maximum male similarity. By inputting the feature amount of the voice data of the registration object and each feature amount of a plurality of pre-stored voice data of a female into the gender recognition model, the similarity between the voice data of the registration object and each of the plurality of voice data of the female is obtained from the gender recognition model, The maximum value among the obtained multiple similarities is calculated as the maximum female similarity. If the maximum male similarity is higher than the maximum female similarity, the gender of the speaker of the registration object voice data is identified as male, If the maximum male similarity is lower than the maximum female similarity, the gender of the speaker of the registration target voice data is identified as female.

8. The speaker recognition method according to claim 5, wherein: In identifying said gender, Calculate the average feature value of a plurality of pre-stored male voice data, The first similarity between the registration object voice data and the male voice data group is obtained from the gender recognition model by inputting the feature value of the registration object voice data and the average feature value of a plurality of male voice data into the gender recognition model. Calculate the average feature value of a plurality of pre-stored female voice data, The second similarity between the registration object voice data and the female voice data group is obtained from the gender recognition model by inputting the feature value of the registration object voice data and the average feature value of a plurality of female voice data into the gender recognition model. If the first similarity is higher than the second similarity, the gender of the speaker of the registration target voice data is identified as male, When the first similarity is lower than the second similarity, the gender of the speaker of the registration target voice data is identified as female.

9. A speaker recognition device, characterized in that include: The recognition target voice data acquisition unit is used to acquire the recognition target voice data; a registration voice data acquisition unit, configured to acquire pre-registered registration voice data; a first extraction unit configured to extract a feature value of the speech data to be recognized; a second extraction unit configured to extract a feature value of the enrollment voice data; a speaker recognition model selection unit that, when the gender of either the speaker of the recognition target speech data or the speaker of the registration speech data is male, selects a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers; and, when the gender of either the speaker of the recognition target speech data or the speaker of the registration speech data is female, selects a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers; and The speaker recognition unit recognizes the speaker of the recognition target speech data by inputting the feature quantity of the recognition target speech data and the feature quantity of the registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, The speaker recognition unit performs the following operations: obtaining, from either the first speaker recognition model or the second speaker recognition model, a degree of similarity between the recognition target speech data and each of the plurality of enrollment speech data by inputting the feature quantity of the recognition target speech data and each of the feature quantities of the plurality of enrollment speech data into either the selected first speaker recognition model or the second speaker recognition model; Identify the speaker of the registered voice data with the highest acquired similarity as the speaker of the recognition target voice data, During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. The speaker recognition unit subtracts the first threshold from the acquired similarity when the similarity is acquired from the first speaker recognition model, and subtracts the second threshold from the acquired similarity when the similarity is acquired from the second speaker recognition model.

10. A speaker recognition device, characterized in that include: The recognition target voice data acquisition unit is used to acquire the recognition target voice data; a registration voice data acquisition unit, configured to acquire pre-registered registration voice data; a first extraction unit configured to extract a feature value of the speech data to be recognized; a second extraction unit configured to extract a feature value of the enrollment voice data; a speaker recognition model selection unit that, when the gender of either the speaker of the recognition target speech data or the speaker of the registration speech data is male, selects a first speaker recognition model that has been machine-learned using male speech data for recognizing male speakers; and, when the gender of either the speaker of the recognition target speech data or the speaker of the registration speech data is female, selects a second speaker recognition model that has been machine-learned using female speech data for recognizing female speakers; and The speaker recognition unit recognizes the speaker of the recognition target speech data by inputting the feature quantity of the recognition target speech data and the feature quantity of the registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, The plurality of registered voice data are associated with identification information for identifying a speaker of each of the plurality of registered voice data. The speaker recognition device further comprises: an identification information acquisition unit for acquiring identification information for identifying a speaker of the recognition target speech data; The registered voice data acquisition unit acquires the registered voice data corresponding to the identification information that matches the acquired identification information from the plurality of registered voice data. The speaker recognition unit performs the following operations: obtaining a similarity between the recognition target speech data and the registration speech data from either the first speaker recognition model or the second speaker recognition model by inputting the feature quantity of the recognition target speech data and the feature quantity of the registration speech data into either the selected first speaker recognition model or the second speaker recognition model; If the obtained similarity is higher than a threshold, the speaker of the registration voice data is identified as the speaker of the recognition target voice data. During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. The speaker recognition unit subtracts the first threshold from the acquired similarity when the similarity is acquired from the first speaker recognition model, and subtracts the second threshold from the acquired similarity when the similarity is acquired from the second speaker recognition model.

11. A program product comprising a speaker recognition program, characterized in that: The speaker recognition program enables the computer to perform the following functions: Obtaining the voice data of the recognition object; Acquire pre-registered registration voice data; Extracting feature quantities of the speech data of the recognition object; extracting a feature value of the enrollment voice data; If the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is male, a first speaker recognition model that has been machine-learned using male voice data for recognizing male speakers is selected; and if the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is female, a second speaker recognition model that has been machine-learned using female voice data for recognizing female speakers is selected; and The speaker of the recognition target speech data is recognized by inputting the feature quantity of the recognition target speech data and the feature quantity of the enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, When identifying the speaker, the speaker recognition program causes the computer to perform the following functions: by inputting the feature value of the recognition target speech data and each feature value of the plurality of registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model, obtaining the similarity between the recognition target speech data and each of the plurality of registration speech data from any one of the first speaker recognition model and the second speaker recognition model; and identifying the speaker of the registration speech data having the highest obtained similarity as the speaker of the recognition target speech data. During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. When identifying the speaker, the speaker recognition program enables the computer to perform the following functions: when the similarity is obtained from the first speaker recognition model, the first threshold is subtracted from the obtained similarity; when the similarity is obtained from the second speaker recognition model, the second threshold is subtracted from the obtained similarity.

12. A program product comprising a speaker recognition program, characterized in that: The speaker recognition program enables the computer to perform the following functions: Obtaining the voice data of the recognition object; Acquire pre-registered registration voice data; Extracting feature quantities of the speech data of the recognition object; extracting a feature value of the enrollment voice data; If the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is male, a first speaker recognition model that has been machine-learned using male voice data for recognizing male speakers is selected; and if the gender of either the speaker of the recognition target voice data or the speaker of the registration voice data is female, a second speaker recognition model that has been machine-learned using female voice data for recognizing female speakers is selected; and The speaker of the recognition target speech data is recognized by inputting the feature quantity of the recognition target speech data and the feature quantity of the enrollment speech data into any one of the selected first speaker recognition model and the second speaker recognition model, The enrollment voice data includes a plurality of enrollment voice data, The plurality of registered voice data are associated with identification information for identifying a speaker of each of the plurality of registered voice data. The speaker recognition program further enables the computer to perform the following functions: obtain recognition information of the speaker used to recognize the recognition target voice data, When acquiring the registered voice data, the speaker recognition program enables the computer to perform the following functions: acquiring, from the plurality of registered voice data, the registered voice data corresponding to the identification information that is consistent with the acquired identification information; When identifying the speaker, the speaker recognition program enables the computer to perform the following functions: by inputting the feature quantity of the recognition target speech data and the feature quantity of the registration speech data into any one of the selected first speaker recognition model and the second speaker recognition model, the similarity between the recognition target speech data and the registration speech data is obtained from any one of the first speaker recognition model and the second speaker recognition model; when the obtained similarity is higher than a threshold value, the speaker of the registration speech data is identified as the speaker of the recognition target speech data, During machine learning, all combinations of feature quantities of two speech data from a plurality of male speech data are input into the first speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the first speaker recognition model. A first threshold value for the similarity that can distinguish the two speech data of the same speaker and the similarity of the two speech data of different speakers is calculated. During machine learning, all combinations of feature quantities of two speech data from a plurality of speech data of a female are input into the second speaker recognition model, and the similarities of the plurality of combinations of the two speech data are obtained from the second speaker recognition model. A second threshold value is calculated for the similarity that can identify the two speech data of the same speaker and the similarity of the two speech data of different speakers. When identifying the speaker, the speaker recognition program enables the computer to perform the following functions: when the similarity is obtained from the first speaker recognition model, the first threshold is subtracted from the obtained similarity; when the similarity is obtained from the second speaker recognition model, the second threshold is subtracted from the obtained similarity.

Citation Information

Patent Citations

  • Voiceprint authentication processing method and apparatus

    CN105513597A

  • User gender identification method and device, storage medium and electronic equipment

    CN110020167A

  • Speaker recognition device, speaker recognition method, and speaker recognition program

    JP2014048534A

  • KR20190024148A