Speaker identification method, speaker identification device, and speaker identification program

The speaker identification method addresses the challenge of identifying unspecified speakers by performing voice recognition, selecting the closest registered speech content, and calculating similarity, ensuring accurate identification despite content mismatches.

JP7737451B2Active Publication Date: 2025-09-10PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023527593
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-11
Filing Date
2022-05-19
Publication Date
2025-09-10
Estimated Expiration
2042-05-19

AI Technical Summary

Technical Problem

Existing speaker identification technologies struggle to identify unspecified speakers when their speech content does not match the content of pre-registered speakers, or when they utter keywords other than predetermined fixed keywords.

Method used

A speaker identification method that performs voice recognition on input speech data, selects the closest registered speech content, identifies a database corresponding to this content, calculates similarity based on stored features, and outputs an identification result, even if the speech content does not match pre-registered content.

Benefits of technology

Enables accurate identification of unspecified speakers by selecting the closest matching registered speech content and calculating similarity, allowing identification even when speech content differs from pre-registered content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007737451000001
    Figure 0007737451000001
  • Figure 0007737451000002
    Figure 0007737451000002
  • Figure 0007737451000003
    Figure 0007737451000003
Patent Text Reader

Abstract

A speaker identification device (1): subjects input utterance data to voice recognition; selects, as a selected piece of utterance content and from among multiple pieces of registered utterance content determined in advance, the registered piece of utterance content that is the closest to recognized piece of utterance content indicated by the voice recognition results; selects, from among multiple databases (41, 42, …, 4N) corresponding to the multiple pieces of registered utterance content, a database corresponding to the selected piece utterance content; calculates a similarity degree between a feature amount of the input utterance data and a feature amount stored in the selected database (4); identifies an unspecified speaker on the basis of the similarity degree; and outputs an identification result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for identifying an unspecified speaker. [Background technology]

[0002] Patent document 1 discloses a technology that performs speech recognition on the generated content of an input pattern and the generated content of a standard pattern, and based on the obtained generated content information, determines a matching section where the generated content of the input pattern matches that of the standard patterns of multiple pre-registered speakers, determines the degree of difference between the input pattern and the standard pattern in the matching section, and recognizes the speaker who generated the input speech based on the determined degree of difference.

[0003] Non-patent document 1 discloses a technology for identifying unspecified speakers by comparing the voice features of predetermined fixed keywords spoken by multiple registered speakers with the voice features of fixed keywords spoken by unspecified speakers.

[0004] However, in the above-mentioned conventional technology, if the speech of an unspecified speaker does not match the speech content of a pre-registered speaker, it is not possible to identify the unspecified speaker, and therefore further improvement was needed. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 3075250 [Non-patent literature]

[0006] [Non-Patent Document 1] Hiroshi Fujimura, Ning Ding, Daichi Hayakawa and Takehiko Kagoshima “Simultaneous Flexible Keyword Detection and Text-dependent Speaker Recognition for Low-resource Devices” Proceedings of the 9th International Conference on Pattern Recognition Applications and Methods (ICPRAM 2020), pages 297-307 Summary of the Invention

[0007] The present disclosure has been made to solve such problems, and aims to provide a technology that can identify an unspecified speaker even if the content of the speech by the unspecified speaker does not match the content of the speech of a pre-registered speaker.

[0008] A speaker identification method in one aspect of the present disclosure is a speaker identification method in a speaker identification device that identifies unspecified speakers, which method acquires input speech data, which is speech data spoken by an unspecified speaker, performs voice recognition on the input speech data, selects, from a predetermined plurality of registered speech contents, a registered speech content that is closest to the recognized speech content indicated by the result of the voice recognition as a selected speech content, selects a database corresponding to the selected speech content from a plurality of databases corresponding to the plurality of registered speech contents, and each database stores features of the speech data when the registered speaker speaks the registered speech content, calculates the similarity between the features of the input speech data and the features stored in the selected database, identifies the unspecified speaker based on the similarity, and outputs an identification result.

[0009] According to the present disclosure, even if the content of speech by an unspecified speaker does not match the content of speech by a registered speaker who has been registered in advance, the unspecified speaker can be identified. [Brief explanation of the drawings]

[0010] [Figure 1] 1 is a block diagram showing an example of the configuration of a speaker identification device 1 according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a data configuration of a database. [Figure 3] 10 is a flowchart illustrating an example of processing performed by the speaker identification device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0011] (Findings underlying this disclosure) A speaker identification technique is known that acquires speech data from an unspecified speaker to be identified, compares the feature values ​​of the acquired speech data with the feature values ​​of the speech data of each of multiple registered speakers, and identifies which of multiple registered speakers the unspecified speaker corresponds to. It has been found that with such speaker identification techniques, the similarity between the feature values ​​of speech data decreases if the speech content differs even for the same speaker, whereas the similarity increases if the speech content is the same for different speakers. In other words, it has been found that the similarity significantly depends on the speech content.

[0012] The technology of Patent Document 1 is based on the premise that there is a matching section in the standard pattern where the speech content matches the input pattern uttered by an unspecified speaker. Therefore, if an unspecified speaker makes an utterance that does not have such a matching section, there is a problem in that the unspecified speaker cannot be identified.

[0013] The technology of Non-Patent Document 1 is premised on the assumption that unspecified speakers will utter predetermined fixed keywords, and does not assume that unspecified speakers will utter keywords other than the fixed keywords. Therefore, the technology of Non-Patent Document 1 has a problem in that it cannot identify unspecified speakers when they utter keywords other than the fixed keywords.

[0014] The present disclosure has been made to solve such problems, and aims to provide a technology that can identify an unspecified speaker even if the content of the speech by the unspecified speaker does not match the content of the speech of a pre-registered speaker.

[0015] A speaker identification method in one aspect of the present disclosure is a speaker identification method in a speaker identification device, which acquires input speech data, which is speech data spoken by an unspecified speaker, performs voice recognition on the input speech data, selects, from a predetermined plurality of registered speech contents, a registered utterance content that is closest to the recognized speech content indicated by the result of the voice recognition as a selected speech content, selects a database corresponding to the selected speech content from a plurality of databases corresponding to the plurality of registered utterance contents, and each database stores features of the speech data when the registered speaker speaks the registered utterance content, calculates a similarity between the features of the input speech data and the features stored in the selected database, identifies the unspecified speaker based on the similarity, and outputs an identification result.

[0016] According to this configuration, input speech data of an unspecified speaker is subjected to speech recognition, a registered utterance content that is closest to the recognized utterance content indicated by the speech recognition result is selected as a selected utterance content from among a plurality of predetermined registered utterance contents, a database corresponding to the selected utterance content is selected from among a plurality of databases, a similarity between the feature values ​​of the registered speaker stored in the selected database and the feature values ​​of the input utterance data is calculated, and the unspecified speaker is identified based on the calculated similarity. Therefore, even if the utterance content of the unspecified speaker does not match the utterance content of the registered speaker registered in advance, the unspecified speaker can be identified.

[0017] In the above speaker identification method, when selecting the selected utterance content, if there is a registered utterance content among the multiple registered utterance contents that matches the recognized utterance content, the matching registered utterance content may be selected as the selected utterance content.

[0018] According to this configuration, if there is a registered utterance content among the multiple registered utterance contents that matches the recognized utterance content, a database corresponding to the matching registered utterance content is selected, and an unspecified speaker is identified using the features of the registered speaker stored in the selected database, thereby enabling unspecified speakers to be identified with high accuracy.

[0019] In the above speaker identification method, when selecting the selected utterance content, if there is no registered utterance content among the multiple registered utterance contents that matches the recognized utterance content, the closest utterance content may be selected as the selected utterance content.

[0020] According to this configuration, if there is no registered utterance content among the multiple registered utterance contents that matches the recognized utterance content, a database corresponding to the registered utterance content that is closest to the recognized utterance content is selected, and an unspecified speaker is identified using the features of the registered speaker stored in the selected database, thereby enabling unspecified speakers to be identified with high accuracy.

[0021] In the above speaker identification method, the selection of the selected utterance content may involve selecting, from the plurality of registered utterance contents, a registered utterance content that includes all of the sound elements included in the recognized utterance content.

[0022] According to this configuration, the registered utterance content that includes all the sound elements contained in the recognized utterance content is selected as the closest registered utterance content from among multiple registered utterance contents, so that the registered utterance content that is closest to the recognized utterance content can be selected with high accuracy.

[0023] In the above speaker identification method, the selection of the selected utterance content may involve selecting from the plurality of registered utterance contents a registered utterance content whose configuration data indicating the configuration of the sound elements is closest to the sound elements contained in the recognized utterance content.

[0024] According to this configuration, the registered utterance content whose constituent data of the sound element is closest to the sound element contained in the recognized utterance content is selected from among a plurality of registered utterance contents, so that the registered utterance content closest to the recognized utterance content can be selected with high accuracy.

[0025] In the above speaker identification method, the sound elements may be phonemes.

[0026] According to this configuration, since phonemes are used as sound elements, it is possible to accurately select the registration utterance content that is closest to the recognized utterance content.

[0027] In the above speaker identification method, the sound element may be a vowel.

[0028] According to this configuration, since vowels are used as sound elements, it is possible to select the registration utterance content that is closest to the recognized utterance content with high accuracy.

[0029] In the above speaker identification method, the phonemes may be a sequence of phonemes in each section when the phonemes included in the utterance content are divided into n sections (n ​​is an integer of 2 or more).

[0030] According to this configuration, since a sequence of phonemes is used as the sound elements, it is possible to accurately select the registration utterance content that is closest to the recognized utterance content.

[0031] In the above speaker identification method, the constituent data may be defined as a vector in which the positions of all sound elements are pre-assigned, and each of one or more sound elements contained in the recognized utterance content or the registered utterance content is assigned a value according to the number of times it appears.

[0032] According to this configuration, the recognized utterance content or the registered utterance content can be expressed by a vector that expresses the characteristics of the sound elements, which makes it easy to calculate the similarity between the registered utterance content and the recognized utterance content.

[0033] In the above speaker identification method, the value corresponding to the number of occurrences may be defined as the ratio of the number of occurrences of each of the one or more sound elements to the total number of sound elements included in the recognized utterance content or the registered utterance content.

[0034] According to this configuration, the value corresponding to the number of occurrences is determined by the proportion of the number of occurrences of each sound element to the total number of sound elements contained in the recognized utterance content or the registered utterance content, so that the characteristics of the sound elements contained in the utterance content can be accurately expressed using vectors.

[0035] In another aspect of the present disclosure, a speaker identification device includes an acquisition unit that acquires input speech data, which is speech data spoken by an unspecified speaker; a recognition unit that performs voice recognition on the input speech data; a first selection unit that selects, from a predetermined plurality of registered speech contents, a registered speech content that is closest to the recognized speech content indicated by the result of the voice recognition as a selected speech content; a second selection unit that selects, from a plurality of databases corresponding to the plurality of registered speech contents, a database corresponding to the selected speech content; each database stores features of the speech data when the registered speaker speaks the registered speech content, a similarity calculation unit that calculates the similarity between the features of the input speech data and the features stored in the selected database; and an output unit that identifies the unspecified speaker based on the similarity and outputs an identification result.

[0036] According to this configuration, it is possible to provide a speaker identification device that can obtain the same effects as the above-described speaker identification method.

[0037] In yet another aspect of the present disclosure, a speaker identification program is a speaker identification program that causes a computer to function as a speaker identification device, and causes the computer to perform the following processes: acquire input speech data, which is speech data spoken by an unspecified speaker; perform voice recognition on the input speech data; select, from a predetermined plurality of registered speech contents, a registered speech content that is closest to the recognized speech content indicated by the result of the voice recognition as a selected speech content; select, from a plurality of databases corresponding to the plurality of registered speech contents, a database corresponding to the selected speech content; each database stores features of the speech data when the registered speaker speaks the registered speech content; calculate a similarity between the features of the input speech data and the features stored in the selected database; identify the unspecified speaker based on the similarity; and output an identification result.

[0038] According to this configuration, it is possible to provide a speaker identification program that can achieve the same effects as the above-described speaker identification method.

[0039] The present disclosure can also be realized as an information updating system that operates using such a speaker identification program. Needless to say, such a speaker identification program can be distributed via a computer-readable non-transitory recording medium such as a CD-ROM or a communication network such as the Internet.

[0040] Note that each of the embodiments described below represents a specific example of the present disclosure. The numerical values, shapes, components, steps, and step orders shown in the following embodiments are merely examples and are not intended to limit the present disclosure. Furthermore, among the components in the following embodiments, components that are not described in the independent claims that represent the highest concept are described as optional components. Furthermore, in all of the embodiments, the respective contents can be combined.

[0041] (Embodiment 1) FIG. 1 is a block diagram showing an example of the configuration of a speaker identification device 1 according to an embodiment of the present disclosure. The speaker identification device 1 is a device that identifies unspecified speakers based on speech data, which is voice data uttered by the unspecified speakers. The unspecified speakers are speakers who have not been identified by the speaker identification device 1. The speaker identification device 1 is implemented in, for example, a smart speaker. However, this is just one example, and the speaker identification device 1 may be implemented in a portable information processing device such as a smartphone or tablet computer, or in a stationary information processing device such as a desktop personal computer.

[0042] The speaker identification device 1 includes a microphone 2, a processor 3, N (N≧2) databases 41, 42, . . . , 4N, an operation unit 5, and a communication circuit 6. The N databases 41, 42, . . . , 4N are collectively referred to as databases 4.

[0043] The microphone 2 collects a sound signal including the voice uttered by the speaker, and inputs the collected sound signal to the acquisition unit 31.

[0044] The processor 3 is configured by, for example, a central processing unit, and includes an acquisition unit 31, a recognition unit 32, a first selection unit 33, a second selection unit 34, a feature calculation unit 35, a similarity calculation unit 36, and an output unit 37. The acquisition unit 31 to the output unit 37 are realized by the processor executing a speaker identification program that causes a computer to function as the speaker identification device 1. However, this is just one example, and the acquisition unit 31 to the output unit 37 may also be configured by dedicated semiconductor circuits such as ASICs (application specific integrated circuits).

[0045] The acquisition unit 31 acquires input speech data, which is speech data uttered by unspecified speakers, from a sound signal input from the microphone 2. For example, the acquisition unit 31 may acquire the input speech data by detecting a speech section from the input sound signal and calculating acoustic features of the detected speech section. The acoustic features may be, for example, Mel-Frequency Cepstrum Coefficients (MFCC) or spectrograms. The acquisition unit 31 may acquire a sound signal from the microphone 2 in response to a start instruction input from the operation unit 5, and acquire the input speech data from the acquired sound signal.

[0046] The recognition unit 32 performs speech recognition on the input utterance data input from the acquisition unit 31, generates recognized utterance content indicating the recognized utterance content, and inputs the generated recognized utterance content to the first selection unit 33. The recognized utterance content is text data in which the input utterance data is expressed in characters. The recognition unit 32 may generate the recognized utterance content using a known speech recognition method. For example, the recognition unit 32 applies an acoustic model such as a hidden Markov model to acoustic features that make up the utterance data to identify phonemes that make up the utterance data, applies a pronunciation dictionary to the identified phonemes to identify words that make up the utterance data, and applies a language model such as an N-gram model to the identified words to generate the utterance content.

[0047] The first selection unit 33 selects, from a plurality of predetermined registered utterance contents, a registered utterance content that is closest to the recognized utterance content input from the recognition unit 32 as a selected utterance content. Databases 41, 42, ..., 4N exist corresponding to N registered utterance contents. The plurality of registered utterance contents are these N registered utterance contents. The registered utterance contents are, for example, commands for the device 100. Examples of commands include "Turn on the TV" to turn on the power of a television, "Turn on the light" to turn on the power of a lighting device, and "Open the window" to open a window of a mobile object or a house.

[0048] Here, if there is a registered utterance content that matches the recognized utterance content among the N registered utterance contents, the first selection unit 33 selects the matching registered utterance content as the selected utterance content. For example, if the recognized utterance content is "Turn on the TV" and the registered utterance contents are "Turn on the TV," "Turn on the lights," and "Open the window," the first selection unit 33 selects "Turn on the TV" as the selected utterance content.

[0049] On the other hand, when there is no registered utterance content that matches the recognized utterance content, the first selection unit 33 may select, as the selected utterance content, the registered utterance content that is closest to the recognized utterance content from among the multiple registered utterance contents.

[0050] For example, the first selection unit 33 may select a registered utterance content that includes all of the sound elements included in the recognized utterance content as the closest registered utterance content. Alternatively, the first selection unit 33 may select, from among a plurality of registered utterance contents, a registered utterance content whose configuration data indicating the configuration of the sound elements is most similar to the sound elements included in the recognized utterance content as the closest registered utterance content.

[0051] A sound element is, for example, a phoneme, a vowel, or a sequence of phonemes. The configuration data is defined by a vector in which the positions of all sound elements are assigned in advance, and each of one or more sound elements included in the recognized utterance content or the registered utterance content is assigned a value according to the number of times it appears. The value according to the number of times it appears is defined, for example, by the ratio of the number of times each of one or more sound elements appears to the total number of sound elements included in the recognized utterance content or the registered utterance content (hereinafter referred to as the "appearance ratio").

[0052] Below, we will explain a specific example in which the closest registered utterance content is selected, using the example where the recognized utterance content is "Turn off the lights" and the registered utterance contents are "Turn on the TV," "Turn on the lights," and "Open the window."

[0053] (Case C1) The phonetic element is a phoneme Phonemes are represented by the 26 letters of the alphabet, from a to z. Therefore, the configuration data indicating the configuration of the phonemes contained in the utterance content (hereinafter referred to as "phoneme configuration data") can be defined as a one-dimensional vector in which each phoneme contained in the utterance content is assigned by its occurrence rate to an array in which phoneme positions are pre-assigned, such as phoneme "a" in the first position, phoneme "b" in the second position, phoneme "z" in the 26th position, and so on.

[0054] For example, "Turn on the TV" can be expressed as "terebitukete" in phonemes, so the phoneme configuration data for "Turn on the TV" is defined as "0, 1 / 12, 0, 0, 4 / 12, . . . , 0".

[0055] This is for the following reason. "terebitukete" is made up of 12 phonemes, so the total number of phonemes is "12." Furthermore, out of the total number "12," the phoneme "b" appears "1 time." Therefore, the occurrence rate of the phoneme "b" is "1 / 12." Similarly, the phoneme "e" appears "4 times," so the occurrence rate of the phoneme "e" is "4 / 12." In phoneme configuration data, the occurrence rate of a phoneme that appears 0 times is "0." From the above, the phoneme configuration data for "Turn on the TV" is "0, 1 / 12, 0, 0, 4 / 12, ..., 0."

[0056] "Turn on the lights" can be expressed as "syoumeitukete" using phonemes, so the phoneme configuration data is "0, 0, 0, 0, 3 / 13, ..., 0". "Open the window" can be expressed as "madowoakete" using phonemes, so the phoneme configuration data is "2 / 11, 0, 0, 1 / 11, 2 / 11, ..., 0". "Turn off the lights" can be expressed as "syoumeikesite" using phonemes, so the phoneme configuration data is "0, 0, 0, 0, 3 / 13, ..., 0".

[0057] The first selection unit 33 calculates the distance between each of the phoneme configuration data of the multiple registered utterance contents and the phoneme configuration data of the recognized utterance content, calculates the similarity so that the shorter the distance, the larger the value, and selects the registered utterance content with the largest calculated similarity as the registered utterance content that is closest to the recognized utterance content.

[0058] The distance is, for example, Euclidean distance. As the similarity, for example, cosine similarity may be adopted. If the constituent data of the registered utterance content is vector v and the constituent data of the recognized utterance content is vector v', the Euclidean distance between vector v and vector v' is D(v, v') = |vv'| 2 The cosine similarity between vector v and vector v' is expressed as Σvi*vi', where i is an index that identifies a phoneme.

[0059] (Case C2) Case where the phonetic element is a vowel The vowels contained in "Turn off the lights" are "i, u, e, o." On the other hand, the vowels contained in "Turn on the TV" are "i, u, e," the vowels contained in "Turn on the lights" are "i, u, e, o," and the vowels contained in "Open the window" are "a, e, o." Of these, the registered utterance content that contains all the vowels of the recognized utterance content is "Turn on the lights," which contains "i, u, e, o." Therefore, the first selection unit 33 selects "Turn on the lights" as the closest utterance content.

[0060] In addition, if there is no registered utterance content that includes all of the vowels contained in the recognized utterance content, the first selection unit 33 may select, as the selected utterance content, the registered utterance content that has the closest structure to the vowels contained in the recognized utterance content.

[0061] Vowels are represented by five letters: a, i, u, e, and o. Therefore, the configuration data indicating the configuration of vowels contained in the utterance (hereinafter referred to as "vowel configuration data") can be defined as a one-dimensional vector in which each vowel contained in the utterance is assigned by its occurrence rate to an array in which vowel positions are pre-assigned, such as vowel "a" in the first position, vowels "i", ..., and vowel "o" in the fifth position. In this case, the occurrence rate is expressed, for example, as the number of times each vowel appears relative to the total number of vowels contained in the recognized utterance or the registered utterance.

[0062] Then, as in case C1, the first selection unit 33 calculates the similarity between the vowel composition data of multiple registered utterance contents and the vowel composition data of the recognized utterance content, and selects the registered utterance content with the greatest similarity as the registered utterance content that is closest to the recognized utterance content.

[0063] (Case C3) The phonetic structure is a sequence of phonemes A phoneme sequence refers to the sequence of phonemes in each section when the phonemes contained in the utterance are divided into n (n ≥ 2) sections. The recognized utterance "Turn off the lights" can be expressed in phonemes as "syoumeikesite." When n = 3, this phoneme is divided into five sections: "syo," "ume," "ike," "sit," and "e." Here, the fifth section has fewer than three phonemes, so it is discarded. Therefore, when n = 3, the phoneme sequence of "Turn off the lights" consists of four elements: "syo," "ume," "ike," and "sit." Hereinafter, these elements will be referred to as "sequence elements."

[0064] Therefore, the phoneme sequence configuration data (hereinafter referred to as "sequence configuration data") can be defined as a one-dimensional vector in which the sequence elements "syo", "ume", "ike", and "sit" are assigned to the 1st, 2nd, 3rd, and 4th positions, respectively, and each sequence element included in the utterance content is assigned by the number of times it appears. Here, in the sequence configuration data, the order of the sequence elements is the order in which the sequence elements appear in the recognized utterance content, but this is just an example, and any order may be adopted.

[0065] When n=3, the sequence elements of "Turn on the lights" are "syo," "ume," "itu," and "ket," and of these, the sequence elements that match the sequence elements of the recognized speech content are "syo" and "ume," and each appears once. Therefore, the sequence configuration data of "Turn on the lights" is "1, 1, 0, 0." The sequence elements of "Turn on the TV" are "ter," "ebi," "tuk," and "ete," and of these, none of the sequence elements match the sequence elements of the recognized speech content. Therefore, the sequence configuration data of "Turn on the TV" is "0, 0, 0, 0." The sequence elements of "Open the window" are "mad," "owo," and "ake," and of these, none of the sequence elements match the sequence elements of the recognized speech content. Therefore, the sequence configuration data of "Open the window" is "0, 0, 0, 0."

[0066] Hereinafter, as in case C1, the first selection unit 33 calculates the similarity between the sequence data of the recognized utterance content and the sequence data of each of the multiple registered utterance contents, and selects the registered utterance content with the greatest similarity as the utterance content closest to the recognized utterance content.

[0067] The second selection unit 34 selects a database 4 corresponding to the selected utterance content input from the first selection unit 33 from the databases 41, 42, . . . , 4N.

[0068] For example, if the registered utterance content corresponding to database 41 is "Turn on the TV," the registered utterance content corresponding to database 42 is "Turn on the lights," the registered utterance content corresponding to database 43 is "Open the window," and the selected utterance content is "Turn on the TV," database 41 is selected.

[0069] FIG. 2 is a diagram showing an example of the data configuration of database 4. Database 4 stores speaker IDs and speaker features (an example of features) in association with each other. A speaker ID is an identifier of a registered speaker. A registered speaker is a speaker whose speaker features are registered in database 4. A registered speaker corresponds to, for example, a person associated with a facility or a mobile object to which the speaker identification device 1 is applied. A facility is, for example, a house, an office, or a school. A person associated with a facility is, for example, a resident of a house, an office staff member, or a school staff member and student. A mobile object is, for example, a passenger car, a bus, a taxi, etc. A person associated with a mobile object is, for example, a driver operating the mobile object.

[0070] The speaker features are features of speech data when a registered speaker utters a registered utterance content. The speaker features are features suitable for speech recognition, such as an i-vector, an x-vector, or a d-vector. In this example, the database 4 stores speaker features of three registered speakers residing in a house. For example, if the database 4 in FIG. 2 corresponds to the registered utterance content "Turn on the TV," the database 4 stores speaker features of speech data when registered speakers U1, U2, and U3 each utter "Turn on the TV." These speaker features are registered in advance in the speaker registration phase.

[0071] In the speaker registration phase, the speaker identification device 1 has each of the registration speakers U1, U2, and U3 speak a plurality of registration utterance contents, collects the spoken sound signals with the microphone 2, acquires speech data from the collected sound signals, calculates speaker features of the acquired speech data, and registers the calculated speaker features in the database 4. Then, when the speaker registration phase ends, the speaker identification device 1 starts the speaker identification phase.

[0072] Returning to FIG. 1, the feature calculation unit 35 calculates speaker features of the input utterance data input from the acquisition unit 31. The configuration of these speaker features is the same as the speaker features registered in the database 4. The feature calculation unit 35 calculates the speaker features using a trained model obtained by machine learning training data in which the input data is utterance data and the output data is a speaker ID. This trained model is composed of a feature extraction unit of the training model, which includes a feature extraction unit and a speaker identification unit. The feature extraction unit extracts speaker features of the input utterance data and inputs the extracted speaker features to the speaker identification unit. The speaker identification unit outputs a speaker ID corresponding to the input speaker features. In the training phase, when utterance data is input to the feature extraction unit, the feature extraction unit and the identification unit are machine-learned so that a speaker ID corresponding to the utterance data is output as the identification result of the identification unit. In the operation phase, the feature extraction unit trained in this way is used as the trained model.

[0073] The similarity calculation unit 36 ​​calculates the similarity between the speaker features of the input utterance data input from the feature calculation unit 35 and the speaker features of each registered speaker stored in the database 4 selected by the second selection unit 34. The similarity has a higher value as the distance between the speaker features of the input utterance data and the speaker features of each registered speaker becomes shorter. The distance is, for example, the Euclidean distance. The similarity may also be a cosine similarity.

[0074] The output unit 37 identifies the unspecified speaker based on the similarity and outputs the identification result to the device 100 using the communication circuit 6. For example, the output unit 37 may identify the registered speaker having the greatest similarity between the speaker feature of the input utterance data and the speaker feature of each registered speaker as the unspecified speaker, generate output data including the identification result, and output the generated output data to the device 100 using the communication circuit 6. For example, the output unit 37 may include the speaker ID of the identified registered speaker in the output data as the identification result. Furthermore, the output data may include the registered utterance content selected by the first selection unit 33 or an identifier that identifies the registered utterance content.

[0075] The operation unit 5 is an input device such as a touch panel, a mouse, a keyboard, buttons, etc. The operation unit 5 receives, for example, an operation from the speaker to instruct the start of speech.

[0076] The device 100 is a device installed in a facility or a mobile object, and is a device that can be connected to and communicate with the speaker identification device 1. When the speaker identification device 1 is installed in a facility, the device 100 is, for example, an electrical device installed in the facility. Examples of electrical devices include air conditioners, televisions, lighting equipment, power windows, power shutters, power curtains, washing machines, refrigerators, and microwave ovens. When the speaker identification device 1 is installed in a mobile object, the device 100 is, for example, a car navigation device, a car air conditioner, a car audio system, wipers, power windows, a control device that controls the drive system of the mobile object, and the like.

[0077] The device 100 and the speaker identification device 1 may be connected via a local area network such as a wireless LAN (Local Area Network), a wired LAN, a CAN (Controller Area Network), etc. When the speaker identification device 1 is configured as a cloud server, the device 100 and the speaker identification device 1 are connected via a wide area network such as the Internet.

[0078] The above is the configuration of the speaker identification device 1. Next, the processing of the speaker identification device 1 will be described. Fig. 3 is a flowchart showing an example of the processing of the speaker identification device 1 in this embodiment. This flowchart starts when, for example, an unspecified speaker inputs an operation to the operation unit 5 to instruct the start of speaking.

[0079] In step S1, the microphone 2 collects a sound signal representing a voice uttered by an unspecified speaker. In step S2, the acquisition unit 31 acquires input speech data by calculating acoustic features of a speech section of the sound signal collected in step S1. As a result, input speech data is acquired in which a sound signal such as "Turn off the lights" is represented by acoustic features.

[0080] In step S3, the recognition unit 32 generates a recognized utterance content by performing voice recognition on the input utterance data, thereby generating a recognized utterance content in which the input utterance data is converted into text data.

[0081] In step S4, the first selection unit 33 determines whether or not there is a registered utterance content that matches the recognized utterance content. In this case, the first selection unit 33 may determine whether or not there is a match by comparing the text data of the recognized utterance content with the text data of the registered utterance content.

[0082] If a matching registered utterance content exists (YES in step S4), the first selection unit 33 selects the matching registered utterance content as the selected utterance content (step S5), and the process proceeds to step S7.

[0083] On the other hand, if there is no matching registered utterance content (NO in step S4), the first selection unit 33 selects, from the multiple registered utterance contents, the registered utterance content that is closest to the recognized utterance content as the selected utterance content (step S6). For example, as described above, the first selection unit 33 may select the registered utterance content that is closest to the recognized utterance content using a method that uses any of phonemes, vowels, or phoneme sequences that make up the recognized utterance content as sound elements. As a result, for example, if there is no registered utterance content that matches the recognized utterance content "Turn off the lights," the registered utterance content that is closest to "Turn off the lights" is selected as the selected utterance content.

[0084] In step S7, the second selection unit 34 selects the database 4 corresponding to the selected utterance content from the databases 41, 42, . . . , 4N.

[0085] In step S8, the feature calculation unit 35 inputs the input utterance data acquired in step S1 into the trained model, and calculates speaker features of the input utterance data.

[0086] In step S9, the similarity calculation unit 36 ​​calculates the similarity between the speaker features of the input utterance data and the speaker features of each registered speaker stored in the database 4 selected in step S7. For example, if there are three registered speakers registered in the selected database 4, the similarity for each of the three speakers is calculated.

[0087] In step S10, the output unit 37 identifies the registered speaker with the greatest similarity among the similarities calculated in step S9 as an unspecified speaker. For example, if the registered speaker U1 has the greatest similarity among the registered speakers U1, U2, and U3, the registered speaker U1 is identified as an unspecified speaker.

[0088] In step S11, the output unit 37 generates output data including a speaker ID indicating the identification result and the registered utterance content, and transmits the generated output data to the device 100 using the communication circuit 6.

[0089] In this way, the speaker identification device 1 performs speech recognition on input utterance data of an unspecified speaker, selects from a predetermined plurality of registered utterance contents the registered utterance content that is closest to the recognized utterance content indicated by the speech recognition result as the selected utterance content, selects a database 4 corresponding to the selected utterance content from databases 41, 42, ..., 4N, calculates the similarity between the speaker features of the registered speaker stored in the selected database 4 and the speaker features of the input utterance data, and identifies the unspecified speaker based on the calculated similarity. Therefore, even if the speech content of the unspecified speaker does not match the speech content of the registered speaker registered in advance, the unspecified speaker can be identified.

[0090] The speaker identification device 1 can be used in the following ways: In one example of a use case, a mobile object is controlled by receiving only commands uttered by the driver. This prevents the mobile object from being controlled by commands uttered by anyone other than the driver, ensuring the safety of the mobile object.

[0091] Another example use case is when a person in a house uses voice to control the device 100 installed in the house. In this case, the device 100 determines the preferences of the person who uttered the command from the input history of the person, and operates in a control mode and a user interface that match the determined preferences.

[0092] The present disclosure can employ the following modifications.

[0093] (1) In the above-mentioned case C2, if there are multiple registered utterance contents that include all of the vowels contained in the recognized utterance content, the first selection unit 33 may select the registered utterance content that has the greatest similarity between the vowel composition data of the recognized utterance content and the vowel composition data of the registered utterance content as the registered utterance content that is closest to the recognized utterance content.

[0094] (2) In the above embodiment, the sound element was one of a phoneme, a vowel, and a sequence of phonemes, but the present disclosure is not limited to this, and the closest registered utterance content may be selected by combining these sound elements.

[0095] For example, the first selection unit 33 may calculate the similarity between the recognized utterance content and each registered utterance content for each phoneme, vowel, and phoneme sequence, calculate the total similarity for each registered utterance content by adding the calculated similarities for each registered utterance content, and select the registered utterance content with the largest total similarity as the closest registered utterance content.

[0096] Alternatively, if the first selection unit 33 cannot uniquely identify the registered utterance content that is closest to the recognized utterance content using the vowels, the first selection unit 33 may select the closest registered utterance content using the phoneme configuration data or the phoneme sequence configuration data. A unique identification is not possible when, for example, there is no registered utterance content that includes all of the vowels included in the recognized utterance content, or when there are multiple registered utterance contents that include all of the vowels included in the recognized utterance content.

[0097] (3) In the above embodiment, when the sound element is a phoneme or a sequence of phonemes, the first selection unit 33 selects the registered utterance content that is closest to the recognized utterance content using the configuration data. However, this is just one example. For example, the first selection unit 33 may select the registered utterance content that includes all of the phonemes or phoneme sequences included in the recognized utterance content as the closest registered utterance content. In this case, if the first selection unit 33 cannot uniquely select the registered utterance content that includes all of the phonemes or phoneme sequences, it can uniquely identify the registered utterance content using the phoneme configuration data or phoneme sequence configuration data described above.

[0098] (4) Some of the blocks constituting the processor 3 and the database 4 may be held by a cloud server.

[0099] (5) The speaker identification device 1 may be implemented in the device 100. [Industrial Applicability]

[0100] The present disclosure is useful in the technical field of identifying speakers by voice.

Claims

1. A speaker identification method in a speaker identification device, comprising: Acquire input utterance data, which is utterance data uttered by an unspecified speaker; performing speech recognition on the input speech data; selecting, as a selected utterance content, a registered utterance content that is closest to the recognized utterance content indicated by the result of the voice recognition from among a plurality of predetermined registered utterance contents; selecting a database corresponding to the selected utterance content from among a plurality of databases corresponding to the plurality of registered utterance contents, and each database stores a feature amount of the utterance data when the registered speaker utters the registered utterance content; Calculating a similarity between the feature of the input speech data and the feature stored in the selected database; identifying the unspecified speaker based on the similarity and outputting the identification result; Speaker identification methods.

2. In selecting the selected utterance content, if there is a registered utterance content that matches the recognized utterance content among the plurality of registered utterance contents, the matching registered utterance content is selected as the selected utterance content.

2. The speaker identification method according to claim 1.

3. In selecting the selected utterance content, if there is no registered utterance content that matches the recognized utterance content among the plurality of registered utterance contents, the closest utterance content is selected as the selected utterance content.

3. The speaker identification method according to claim 1 or 2.

4. In the selection of the selected utterance content, a registered utterance content including all of the sound elements included in the recognized utterance content is selected from the plurality of registered utterance contents.

2. The speaker identification method according to claim 1.

5. In the selection of the selected utterance content, a registered utterance content having configuration data indicating a configuration of a sound element that is closest to a sound element included in the recognized utterance content is selected from the plurality of registered utterance contents.

2. The speaker identification method according to claim 1.

6. The phonetic elements are phonemes.

6. The speaker identification method according to claim 4 or 5.

7. The sound element is a vowel.

6. The speaker identification method according to claim 4 or 5.

8. The phonetic elements are phoneme sequences in each section when the phonemes included in the utterance content are divided into n sections (n ​​is an integer of 2 or more), 6. The speaker identification method according to claim 4 or 5.

9. the configuration data is defined by a vector in which the positions of all sound elements are assigned in advance, and each of one or more sound elements included in the recognized utterance content or the registered utterance content is assigned a value according to the number of times of appearance.

6. The speaker identification method according to claim 5.

10. the value according to the number of occurrences is defined as a ratio of the number of occurrences of each of the one or more sound elements to the total number of sound elements included in the recognized utterance content or the registered utterance content; 10. The speaker identification method of claim 9.

11. an acquisition unit that acquires input utterance data that is utterance data uttered by an unspecified speaker; a recognition unit that performs voice recognition on the input speech data; a first selection unit that selects, as a selected utterance content, a registered utterance content that is closest to a recognized utterance content indicated by the result of the speech recognition from a plurality of predetermined registered utterance contents; a second selection unit that selects a database corresponding to the selected utterance content from a plurality of databases corresponding to the plurality of registered utterance contents, and each database stores a feature amount of the utterance data when a registered speaker utters the registered utterance content; a similarity calculation unit that calculates a similarity between a feature of the input speech data and a feature stored in a selected database; an output unit that identifies the unspecified speaker based on the similarity and outputs the identification result. Speaker identification device.

12. A speaker identification program that causes a computer to function as a speaker identification device, The computer, Acquire input utterance data, which is utterance data uttered by an unspecified speaker; performing speech recognition on the input speech data; selecting, as a selected utterance content, a registered utterance content that matches or is closest to the recognized utterance content indicated by the result of the voice recognition from among a plurality of predetermined registered utterance contents; selecting a database corresponding to the selected utterance content from among a plurality of databases corresponding to the plurality of registered utterance contents, and each database stores a feature amount of the utterance data when the registered speaker utters the registered utterance content; Calculating a similarity between the feature of the input speech data and the feature stored in the selected database; Execute a process to identify the unspecified speaker based on the similarity and output the identification result. Speaker identification program.

Citation Information

Patent Citations

  • Control method of voice recognition device

    JP2004301893A

  • Voice recognizer

    JP2009145755A

  • Speaker recognition method and device

    JP3075250B2