Information processing method, information processing device, and information processing program

By converting enrollment voices into multiple acoustic variants and calculating a threshold based on feature comparisons, the method enhances speaker recognition accuracy despite changes in voice characteristics.

JP7792430B2Active Publication Date: 2025-12-25PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2023557631
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-08
Filing Date
2022-08-23
Publication Date
2025-12-25
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

Existing speaker recognition technologies struggle to maintain performance due to changes in acoustic characteristics of a speaker's voice, such as those caused by physical conditions like a cold, leading to misidentification.

Method used

The method involves converting an enrollment voice into multiple characteristic-converted voices with different acoustic characteristics, extracting speaker features, comparing all combinations of these features, and calculating a threshold for accurate speaker recognition based on these comparisons.

Benefits of technology

This approach reduces the degradation of speaker recognition performance by determining whether the similarity between input and enrollment voices falls within an acceptable range despite changes in acoustic characteristics, ensuring accurate speaker identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007792430000001
    Figure 0007792430000001
  • Figure 0007792430000002
    Figure 0007792430000002
  • Figure 0007792430000003
    Figure 0007792430000003
Patent Text Reader

Abstract

In the present invention, a speaker recognition device acquires registered speech, converts the acquired registered speech to a plurality of characteristic converted speech items having different acoustic characteristics, extracts a speaker feature amount indicating the feature of a speaker from the registered speech, extracts speaker feature amounts from the plurality of characteristic converted speech items, makes a comparison for all combinations of two speaker feature amounts among some or all of the speaker feature amount extracted from the registered speech and the plurality of speaker feature amounts extracted from the plurality of characteristic converted speech items, and calculates a threshold value for use in recognizing the speaker of input speech on the basis of the comparison result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a technique for recognizing the speaker of an input voice by comparing the input voice with an enrollment voice. [Background technology]

[0002] For example, in the speech recognition device shown in Patent Document 1, multiple variant speeches with individually different characteristics for a single speech are registered in multiple speech recognition circuits, the speech input by the speaker is compared with each of the multiple registered variant speeches, and the input speech is recognized from each comparison result.

[0003] Furthermore, for example, in the speaker identification model training method shown in Patent Document 2, second voice data of a second speaker is generated by performing a voice quality conversion process on first voice data of a first speaker, and the speaker identification model training process is performed using the first voice data and the second voice data as training data.

[0004] However, with the above-mentioned conventional techniques, it is difficult to reduce the deterioration of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input voice, and further improvement is needed. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Publication No. 124599 / 1983 [Patent Document 2] Patent Publication No. 2021-33260 Summary of the Invention

[0006] The present disclosure has been made to solve the above problems, and aims to provide a technology that can reduce degradation of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input voice.

[0007] The information processing method according to the present disclosure includes a computer acquiring a registration voice, converting the registration voice into a plurality of characteristic-converted voices each having different acoustic characteristics, extracting speaker features indicating speaker characteristics from the registration voice, extracting the speaker features from each of the plurality of characteristic-converted voices, comparing all combinations of two speaker features, among some or all of the speaker features extracted from the registration voice and the plurality of speaker features extracted from the plurality of characteristic-converted voices, and calculating a threshold value to be used for recognizing the speaker of the input voice based on the comparison results.

[0008] According to the present disclosure, it is possible to reduce the degradation of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input voice. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram illustrating a configuration of a speaker recognition device according to an embodiment of the present disclosure. [Figure 2] FIG. 10 is a diagram showing an example of similarities of all combinations of two speaker features among a plurality of speaker features in the present embodiment. [Figure 3] 10 is a flowchart illustrating the operation of a voice registration process of the speaker recognition device according to the present embodiment. [Figure 4] FIG. 10 is a diagram showing an example of a cumulative distribution function F(x) calculated based on a plurality of similarities. [Figure 5] 10 is a flowchart for explaining the operation of a speaker recognition process of the speaker recognition device according to the present embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] (Findings that formed the basis of this disclosure) Even when the voice is spoken by the same person, voice characteristics can change due to changes in physical condition, such as catching a cold. When voice characteristics change due to changes in physical condition, speaker recognition that recognizes a speaker from a person's voice may recognize a different speaker even though the voice is the same person as a pre-registered voice. To solve this problem, a speaker recognition method that takes into account changes in voice due to changes in physical condition is needed.

[0011] In the speech recognition device disclosed in Patent Document 1, when speech registration is performed, a speech signal of a desired word to be recognized is input from a microphone, and the input speech signal is branched and introduced into each of a plurality of characteristic variable circuits. The plurality of characteristic variable circuits then perform speech transformation according to their amplification degrees and bandpass characteristics, and the created plurality of transformed speech signals are stored in each of a plurality of speech recognition circuits as a plurality of registered speeches. When speech recognition is performed, a speech signal is input from the microphone, and the input speech signal is branched and passes directly through each of a plurality of characteristic variable circuits before being introduced into each of a plurality of speech recognition circuits. Each of the plurality of speech recognition circuits then compares the speech pattern of the input speech with that of the stored registered speech, and speech recognition processing is performed based on the comparison results of each speech recognition circuit.

[0012] However, this speech recognition device has difficulty recognizing input speech when a change that cannot be represented by the characteristic variable circuit occurs in the input speech. Furthermore, the speech recognition device of Patent Document 1 is an invention related to speech recognition, which recognizes speech, and does not take into consideration determining the identity of a speaker in speaker recognition, which recognizes a speaker. Therefore, the technology of Patent Document 1 has difficulty reducing degradation of speaker recognition performance due to changes in the acoustic characteristics of a speaker's input speech.

[0013] Furthermore, in the speaker identification model training process disclosed in Patent Document 2, in order to expand the training data of the speaker identification model, a voice conversion process is performed on first voice data of a first speaker to generate second voice data of a second speaker different from the first speaker. Then, the speaker identification model training process is performed using the first voice data and the second voice data as training data. Furthermore, in the speaker identification process, voice data is input to a speaker identification model that has been trained in advance, and speaker identification information is output from the speaker identification model.

[0014] In the above-mentioned Patent Document 2, a speaker identification model is trained using first speech data of a first speaker and second speech data of a second speaker. Therefore, even if speech data of the first speaker with changed acoustic characteristics is input to the speaker identification model, the speaker identification model cannot determine that the speaker of the speech data is the first speaker, and may determine that the speaker is not the first speaker but a different speaker. Therefore, with the technology of Patent Document 2, it is difficult to reduce degradation of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input speech.

[0015] In order to solve the above problems, the following techniques are disclosed.

[0016] (1) An information processing method according to one aspect of the present disclosure includes a computer acquiring a registration voice, converting the registration voice into a plurality of characteristic-converted voices each having different acoustic characteristics, extracting speaker features indicating speaker characteristics from the registration voice, extracting the speaker features from each of the plurality of characteristic-converted voices, comparing all combinations of two speaker features, among some or all of the speaker features extracted from the registration voice and the plurality of speaker features extracted from the plurality of characteristic-converted voices, and calculating a threshold value to be used for recognizing the speaker of the input voice based on the comparison results.

[0017] According to this configuration, the registered voice is converted into a plurality of characteristic-converted voices each having different acoustic characteristics, and all combinations of two speaker features, among some or all of the speaker features extracted from the registered voice and the plurality of speaker features extracted from the plurality of characteristic-converted voices, are compared, and a threshold value used to recognize the speaker of the input voice is calculated based on the comparison results.

[0018] Therefore, even if the similarity between the input voice and the registered voice decreases due to a change in the acoustic characteristics of the speaker's input voice, it is possible to determine using a threshold whether the decrease in similarity is within an acceptable range depending on the change in acoustic characteristics, thereby reducing the deterioration of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input voice.

[0019] (2) In the information processing method described in (1) above, the threshold value may be stored in a memory in association with the registered voice, an input voice of a speaker to be recognized may be acquired, the registered voice may be acquired, the speaker features may be extracted from the input voice, the speaker features may be extracted from the registered voice, a similarity between the speaker features extracted from the input voice and the speaker features extracted from the registered voice may be calculated, and if the calculated similarity is greater than the threshold value stored in the memory, a recognition result may be output indicating that the speaker of the input voice matches the speaker of the registered voice.

[0020] According to this configuration, the similarity between the speaker features extracted from the input speech and the speaker features extracted from the enrollment speech is calculated, and the calculated similarity is compared with a pre-stored threshold. The threshold is calculated based on the comparison results of all combinations of two speaker features, among some or all of the speaker features extracted from the enrollment speech and the multiple speaker features extracted from the multiple characteristic-converted speeches. If the calculated similarity is greater than the threshold, a recognition result indicating that the speaker of the input speech matches the speaker of the enrollment speech is output. Therefore, even if the acoustic characteristics of the input speech change, it is possible to accurately recognize whether the speaker of the input speech matches the speaker of the enrollment speech.

[0021] (3) In the information processing method described in (1) or (2) above, in comparing all combinations of the two speaker features, the similarity of all combinations of the two speaker features, among some or all of the speaker features extracted from the registered voice and the plurality of speaker features extracted from the plurality of characteristic-converted voices, may be calculated, and in calculating the threshold, the threshold may be calculated based on the calculated similarities.

[0022] According to this configuration, the threshold can be calculated using the similarity of all combinations of two speaker features, among some or all of the speaker features extracted from the registered voice and the multiple speaker features extracted from the multiple characteristic-converted voices.

[0023] (4) In the information processing method described in (3) above, in calculating the threshold, the minimum similarity among the calculated similarities may be calculated as the threshold.

[0024] According to this configuration, during speaker recognition, the smallest similarity among multiple similarities can be used as a threshold to determine whether the speaker of the input voice whose acoustic characteristics have changed is the same as the speaker of the registered voice.

[0025] (5) In the information processing method described in (3) above, in calculating the threshold, an average of the calculated similarities may be calculated as the threshold.

[0026] According to this configuration, during speaker recognition, the average of multiple similarities can be used as a threshold to determine whether the speaker of the input voice whose acoustic characteristics have changed is the same as the speaker of the registered voice.

[0027] (6) In the information processing method described in (3) above, in calculating the threshold, a cumulative distribution function F(x) indicating the proportion of the calculated similarities whose similarity values ​​are equal to or less than x may be calculated, and the value of the similarity when the calculated cumulative distribution function F(x) is a predetermined proportion may be calculated as the threshold.

[0028] According to this configuration, it is possible to determine whether the speaker of the input voice with changed acoustic characteristics is the same as the speaker of the registered voice by using the similarity value when the cumulative distribution function F(x), which indicates the proportion of similarity values ​​that are less than or equal to x, among the multiple calculated similarities, is a predetermined proportion as a threshold.

[0029] (7) In the information processing method described in any one of (1) to (6) above, the plurality of characteristic-converted voices may include a first characteristic-converted voice in which the voice quality of the registered voice is converted into a voice quality that has changed due to the effects of a cold.

[0030] According to this configuration, the threshold is calculated taking into account the first characteristic converted voice in which the voice quality of the registered voice has been converted into a voice quality that has changed due to the influence of a cold, thereby improving the performance of recognizing whether the speaker of the input voice whose voice quality has changed due to the influence of a cold is the same as the speaker of the registered voice.

[0031] (8) In the information processing method described in any one of (1) to (6) above, the plurality of characteristic-converted voices may include a second characteristic-converted voice obtained by converting the speech content of the registration voice into a different speech content.

[0032] According to this configuration, the threshold is calculated taking into account the second characteristic converted voice in which the speech content of the registered voice is converted into different speech content, thereby improving the performance of recognizing whether the speaker of the input voice, whose speech content is different from the registered voice, is the same as the speaker of the registered voice.

[0033] (9) In the information processing method described in any one of (1) to (6) above, the plurality of characteristic-converted sounds may include a third characteristic-converted sound obtained by adding noise to the registered sound.

[0034] According to this configuration, the threshold is calculated taking into account the third characteristic converted voice in which noise is added to the registered voice, thereby improving the performance of recognizing whether the speaker of the input voice containing noise is the same as the speaker of the registered voice.

[0035] (10) In the information processing method described in any one of (1) to (6) above, the plurality of characteristic-converted voices may include a fourth characteristic-converted voice in which the speaking speed of the registered voice is changed to a different speaking speed.

[0036] According to this configuration, the threshold is calculated taking into account the fourth characteristic converted voice in which the speaking speed of the registered voice is changed to a different speaking speed, thereby improving the performance of recognizing whether the speaker of the input voice, whose speaking speed is different from the registered voice, is the same as the speaker of the registered voice.

[0037] (11) In the information processing method described in any one of (1) to (6) above, the plurality of characteristic-converted voices may include a fifth characteristic-converted voice in which the voice quality of the registered voice is converted into a voice quality that expresses a predetermined emotion.

[0038] According to this configuration, the threshold is calculated taking into account the fifth characteristic converted voice in which the voice quality of the registered voice is converted into a voice quality that expresses a predetermined emotion, thereby improving the performance of recognizing whether the speaker of the input voice spoken in a voice quality that expresses a predetermined emotion is the same as the speaker of the registered voice.

[0039] Furthermore, the present disclosure can be realized not only as an information processing method that executes the characteristic processes described above, but also as an information processing device having a characteristic configuration corresponding to the characteristic processes executed by the information processing method. Furthermore, the present disclosure can also be realized as a computer program that causes a computer to execute the characteristic processes included in such an information processing method. Therefore, the same effects as those of the above information processing method can also be achieved in the following other aspects.

[0040] (12) An information processing device according to another aspect of the present disclosure includes an acquisition unit that acquires a registered voice, a conversion unit that converts the acquired registered voice into a plurality of characteristic-converted voices each having different acoustic characteristics, a first extraction unit that extracts speaker features indicating speaker characteristics from the registered voice, a second extraction unit that extracts the speaker features from each of the plurality of characteristic-converted voices, a comparison unit that compares all combinations of two speaker features, among some or all of the speaker features extracted from the registered voice and a plurality of speaker features extracted from the plurality of characteristic-converted voices, and a calculation unit that calculates a threshold value used to recognize the speaker of the input voice based on the comparison result.

[0041] (13) An information processing program according to another aspect of the present disclosure causes a computer to acquire a registered voice, convert the acquired registered voice into a plurality of characteristic-converted voices each having different acoustic characteristics, extract speaker features indicating speaker characteristics from the registered voice, extract the speaker features from each of the plurality of characteristic-converted voices, compare all combinations of two speaker features, among some or all of the speaker features extracted from the registered voice and the plurality of speaker features extracted from the plurality of characteristic-converted voices, and calculate a threshold value to be used for recognizing the speaker of the input voice based on the comparison results.

[0042] (14) A non-transitory computer-readable recording medium having recorded thereon an information processing program relating to another aspect of the present disclosure causes a computer to acquire a registered voice, convert the acquired registered voice into a plurality of characteristic-converted voices each having different acoustic characteristics, extract speaker features indicating speaker characteristics from the registered voice, extract the speaker features from each of the plurality of characteristic-converted voices, compare all combinations of two speaker features, among the speaker features extracted from the registered voice and some or all of the plurality of speaker features extracted from the plurality of characteristic-converted voices, and calculate a threshold value to be used for recognizing the speaker of the input voice based on the comparison results.

[0043] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Note that the following embodiments are examples of specific embodiments of the present disclosure and do not limit the technical scope of the present disclosure.

[0044] (Embodiment) FIG. 1 is a block diagram showing a configuration of a speaker recognition device 1 according to an embodiment of the present disclosure.

[0045] The speaker recognition device 1 compares the input speech to be recognized with pre-stored registered speech, and if the input speech and the registered speech are similar, determines that the speaker of the input speech is the same as the speaker of the registered speech.

[0046] Speaker recognition includes speaker authentication (or speaker verification) that compares one input voice with one registered voice to determine whether they belong to the same person, and speaker identification that compares one input voice with each of multiple registered voices to determine which of the multiple registered voices the speaker of the input voice is. The speaker recognition device 1 in this embodiment mainly performs speaker authentication, but the present disclosure is not particularly limited to this, and speaker identification may also be performed.

[0047] The speaker recognition device 1 includes a speaker ID acquisition unit 11, a registered voice acquisition unit 12, a characteristic conversion unit 13, a first feature extraction unit 14, a second feature extraction unit 15, a first feature comparison unit 16, a threshold calculation unit 17, a registered voice storage unit 18, an input voice acquisition unit 19, a third feature extraction unit 20, a fourth feature extraction unit 21, a second feature comparison unit 22, a speaker recognition unit 23, and a recognition result output unit 24.

[0048] The speaker ID acquisition unit 11, the registered voice acquisition unit 12, the characteristic conversion unit 13, the first feature extraction unit 14, the second feature extraction unit 15, the first feature comparison unit 16, the threshold calculation unit 17, the input voice acquisition unit 19, the third feature extraction unit 20, the fourth feature extraction unit 21, the second feature comparison unit 22, the speaker recognition unit 23, and the recognition result output unit 24 are realized by a processor. The processor is composed of, for example, a CPU (Central Processing Unit) or the like.

[0049] The registered voice storage unit 18 is realized by a memory, which may be, for example, a read-only memory (ROM) or an electrically erasable programmable read-only memory (EEPROM).

[0050] The speaker recognition device 1 may be, for example, a computer, a smartphone, a tablet computer, or a server. The speaker recognition device 1 may also be incorporated into other devices such as a car navigation device or a home appliance.

[0051] The speaker recognition device 1 executes a voice registration process for registering a speaker's voice in advance, and a speaker recognition process for recognizing a speaker from an input voice. The speaker recognition device 1 can switch between the voice registration process and the speaker recognition process in response to an instruction from a user.

[0052] The speaker ID acquisition unit 11 acquires a speaker ID for recognizing a speaker. In a voice registration process, the speaker ID acquisition unit 11 acquires a speaker ID for recognizing a speaker to be registered. In a speaker recognition process, the speaker ID acquisition unit 11 acquires a speaker ID for recognizing a speaker to be recognized. For example, the speaker ID acquisition unit 11 may be connected to a camera (not shown). The camera may capture an image of a barcode or two-dimensional code in which the speaker ID is stored. The speaker ID acquisition unit 11 may acquire the speaker ID from an image of the barcode or two-dimensional code captured by the camera.

[0053] The speaker ID acquisition unit 11 may be connected to an IC card reader (not shown). The IC card reader may read information stored in an IC card in a contactless manner using NFC (Near Field Communication). The IC card stores a speaker ID in advance. A speaker brings an IC card storing the speaker ID close to the IC card reader. This causes the IC card reader to acquire the speaker ID from the nearby IC card. The speaker ID acquisition unit 11 may acquire the speaker ID from the IC card reader.

[0054] The speaker ID acquiring unit 11 may be connected to an input device such as a keyboard or a touch panel (not shown). The input device may accept input of a speaker ID by a speaker. The speaker ID acquiring unit 11 may acquire the speaker ID input by the speaker from the input device.

[0055] In the voice registration process or the speaker recognition process, the speaker ID acquisition unit 11 outputs the acquired speaker ID to the registration voice acquisition unit 12.

[0056] The registration voice acquisition unit 12 acquires the registration voice. The registration voice acquisition unit 12 may be connected to a microphone (not shown). The microphone collects the voice spoken by the speaker, converts it into voice data, and outputs it to the speaker recognition device 1. In the voice registration process, the microphone outputs the registration voice spoken by the speaker to be registered to the speaker recognition device 1.

[0057] In the voice registration process, the registration voice acquisition unit 12 acquires the registration voice of the speaker to be registered from the microphone. The registration voice acquisition unit 12 stores the acquired registration voice in the registration voice storage unit 18 in association with the speaker ID acquired by the speaker ID acquisition unit 11. Furthermore, in the voice registration process, the registration voice acquisition unit 12 outputs the acquired registration voice to the first feature extraction unit 14 and the characteristic conversion unit 13.

[0058] In addition, if the registration voice of the speaker to be registered is stored in advance in the registration voice storage unit 18, in the voice registration process, the registration voice acquisition unit 12 may acquire the registration voice from the registration voice storage unit 18. In addition, in the voice registration process, the registration voice acquisition unit 12 may acquire the registration voice from other external devices such as a smartphone, a tablet computer, or a personal computer.

[0059] In addition, in the speaker recognition process, the registration voice acquisition unit 12 acquires the registration voice corresponding to the speaker ID acquired by the speaker ID acquisition unit 11 from the registration voice storage unit 18. In the speaker recognition process, the registration voice acquisition unit 12 outputs the acquired registration voice to the fourth feature extraction unit 21.

[0060] The characteristic conversion unit 13 converts the registered voice acquired by the registered voice acquisition unit 12 into a plurality of characteristic-converted voices each having a different acoustic characteristic. The characteristic conversion unit 13 includes a first characteristic conversion unit 131, a second characteristic conversion unit 132, a third characteristic conversion unit 133, a fourth characteristic conversion unit 134, and a fifth characteristic conversion unit 135.

[0061] The first characteristic conversion unit 131 generates first characteristic-converted voice by converting the voice quality of the registered voice into voice quality changed due to the influence of a cold. For example, the first characteristic conversion unit 131 converts the registered voice into first characteristic-converted voice with a husky voice quality. The first characteristic conversion unit 131 is realized by a voice conversion model trained by machine learning so that, when a voice is input, the voice quality of the input voice is converted into voice quality changed due to the influence of a cold and output as first characteristic-converted voice. The voice conversion model is, for example, a neural network model. The voice conversion model is generated by machine learning using, as training data, a voice uttered by a person in a normal state and a voice uttered by a person when the voice quality has changed due to the influence of a cold, with the input being the normal voice and the output being the voice when the voice quality has changed due to the influence of a cold.

[0062] Examples of machine learning include supervised learning, which learns the relationship between input and output using training data in which labels (output information) are assigned to the input information; unsupervised learning, which builds a data structure from only unlabeled input; semi-supervised learning, which handles both labeled and unlabeled data; and reinforcement learning, which learns behaviors that maximize rewards through trial and error. Specific machine learning techniques include neural networks (including deep learning using multi-layer neural networks), genetic programming, decision trees, Bayesian networks, and support vector machines (SVMs). Any of the above specific examples can be used in the machine learning of a voice conversion model.

[0063] The second characteristic conversion unit 132 generates second-characteristic-converted speech by converting the speech content of the registered speech into different speech content. For example, if the speech content of the registered speech is "Good morning," the second characteristic conversion unit 132 converts the speech content of the registered speech into second-characteristic-converted speech with different speech content, such as "Hello." The second characteristic conversion unit 132 is realized by an utterance content conversion model that is machine-learned so that, when speech is input, the speech content of the input speech is converted into different speech content and output as second-characteristic-converted speech. The utterance content conversion model is, for example, a neural network model. The utterance content conversion model is generated by machine learning using speech in which a person speaks a first phrase and speech in which a person speaks a second phrase different from the first phrase as training data, with the input being speech of the first phrase and the output being speech of the second phrase. In the machine learning of the utterance content conversion model, any of the specific examples listed above may be used.

[0064] The third characteristic conversion unit 133 generates third characteristic-converted speech by adding noise to the registered speech. The noise is, for example, environmental noise such as the sound of passing vehicles and the sound of falling rain. The noise is pre-stored in memory within the speaker recognition device 1. The third characteristic conversion unit 133 reads the noise from the memory and adds the read noise to the registered speech to generate third characteristic-converted speech.

[0065] The fourth characteristic conversion unit 134 generates fourth characteristic-converted speech by changing the speaking speed of the registered speech to a different speaking speed. For example, the fourth characteristic conversion unit 134 generates fourth characteristic-converted speech by changing the speaking speed of the registered speech from a first speed to a second speed that is faster than the first speed. The fourth characteristic conversion unit 134 may also generate fourth characteristic-converted speech by changing the speaking speed of the registered speech from the first speed to a third speed that is slower than the first speed.

[0066] The fifth characteristic conversion unit 135 generates fifth characteristic-converted speech by converting the voice quality of the registered speech into a voice quality that expresses a predetermined emotion. For example, the fifth characteristic conversion unit 135 converts the voice quality of the registered speech into a fifth characteristic-converted speech having a voice quality that expresses a happy emotion, a voice quality that expresses an angry emotion, or a voice quality that expresses a sad emotion. The fifth characteristic conversion unit 135 is realized by a voice conversion model that has been machine-learned so that, when a speech is input, the voice quality of the input speech is converted into a voice quality that expresses a happy emotion, a voice quality that expresses an angry emotion, or a voice quality that expresses a sad emotion, and the converted speech is output. The voice conversion model is, for example, a neural network model. The voice conversion model is generated by machine learning using speech uttered by a person in a normal state and speech uttered by a person in a happy, angry, or sad state as training data, with the input being the normal speech and the output being the speech uttered when the person is in a happy, angry, or sad state. In the machine learning of the voice conversion model, any of the specific examples listed above may be used.

[0067] In the present embodiment, five characteristic-converted sounds with different acoustic characteristics are generated by the first characteristic conversion unit 131 to the fifth characteristic conversion unit 135, but the present disclosure is not particularly limited to this, and four or less characteristic-converted sounds from the first characteristic-converted sound to the fifth characteristic-converted sound may be generated. Also, a characteristic-converted sound different from the first characteristic-converted sound to the fifth characteristic-converted sound may be generated, or six or more characteristic-converted sounds may be generated.

[0068] The first feature extraction unit 14 extracts speaker features indicating speaker characteristics from the enrollment speech acquired by the enrollment speech acquisition unit 12. The speaker features are, for example, i-vectors. The i-vectors are low-dimensional vector features extracted from speech data by applying factor analysis to a GMM (Gaussian Mixture Model) supervector. Note that the method for extracting i-vectors is a conventional technique, and therefore a detailed description thereof will be omitted. Furthermore, the speaker features are not limited to i-vectors and may be other features such as x-vectors. The first feature extraction unit 14 outputs the speaker features extracted from the enrollment speech to the first feature comparison unit 16.

[0069] The second feature extraction unit 15 extracts speaker features from each of the plurality of characteristic-converted voices converted by the characteristic conversion unit 13. The speaker features are, for example, i-vectors. The second feature extraction unit 15 includes a fifth feature extraction unit 151, a sixth feature extraction unit 152, a seventh feature extraction unit 153, an eighth feature extraction unit 154, and a ninth feature extraction unit 155.

[0070] The fifth feature extraction unit 151 extracts speaker features from the first characteristic-converted speech converted by the first characteristic conversion unit 131. The fifth feature extraction unit 151 outputs the speaker features extracted from the first characteristic-converted speech to the first feature comparison unit 16.

[0071] The sixth feature extraction unit 152 extracts speaker features from the second characteristic-converted speech converted by the second characteristic conversion unit 132. The sixth feature extraction unit 152 outputs the speaker features extracted from the second characteristic-converted speech to the first feature comparison unit 16.

[0072] The seventh feature extraction unit 153 extracts speaker features from the third characteristic-converted speech converted by the third characteristic conversion unit 133. The seventh feature extraction unit 153 outputs the speaker features extracted from the third characteristic-converted speech to the first feature comparison unit 16.

[0073] The eighth feature extraction unit 154 extracts speaker features from the fourth characteristic converted speech converted by the fourth characteristic conversion unit 134. The eighth feature extraction unit 154 outputs the speaker features extracted from the fourth characteristic converted speech to the first feature comparison unit 16.

[0074] The ninth feature extraction unit 155 extracts speaker features from the fifth characteristic converted speech converted by the fifth characteristic conversion unit 135. The ninth feature extraction unit 155 outputs the speaker features extracted from the fifth characteristic converted speech to the first feature comparison unit 16.

[0075] The first feature comparing unit 16 compares all combinations of two speaker features, among the speaker features extracted from the enrollment speech and some or all of the plurality of speaker features extracted from the plurality of characteristic-converted speeches. More specifically, the first feature comparing unit 16 calculates the similarity between all combinations of two speaker features, among the speaker features extracted from the enrollment speech and some or all of the plurality of speaker features extracted from the plurality of characteristic-converted speeches.

[0076] The first feature comparison unit 16 calculates the similarity using a model based on Probabilistic Linear Discriminant Analysis (PLDA). The PLDA model automatically selects features effective for speaker recognition from 400-dimensional i-vector features, and calculates the log-likelihood ratio as the similarity.

[0077] In this embodiment, the first feature comparison unit 16 calculates the similarity between all combinations of two speaker features, among all speaker features extracted from the registered voice and multiple speaker features extracted from multiple characteristic-converted voices.

[0078] FIG. 2 is a diagram showing an example of the similarity of all combinations of two speaker features among a plurality of speaker features in this embodiment.

[0079] As shown in Figure 2, the similarity between the speaker features of the registered voice and the speaker features of the first characteristic-converted voice is 34.5, the similarity between the speaker features of the registered voice and the speaker features of the second characteristic-converted voice is 40.1, and the similarity between the speaker features of the first characteristic-converted voice and the speaker features of the second characteristic-converted voice is 31.7.

[0080] In addition to the similarities shown in FIG. 2 , the first feature comparison unit 16 also compares the similarities between the speaker features of the registered voice and the speaker features of the third characteristic-converted voice, between the speaker features of the registered voice and the speaker features of the fourth characteristic-converted voice, between the speaker features of the registered voice and the speaker features of the fifth characteristic-converted voice, between the speaker features of the first characteristic-converted voice and the speaker features of the third characteristic-converted voice, between the speaker features of the first characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, and between the speaker features of the first characteristic-converted voice and the speaker features of the fifth characteristic-converted voice. The similarity between the speaker features of the second characteristic-converted voice and the speaker features of the third characteristic-converted voice, the similarity between the speaker features of the second characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, the similarity between the speaker features of the second characteristic-converted voice and the speaker features of the fifth characteristic-converted voice, the similarity between the speaker features of the third characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, the similarity between the speaker features of the third characteristic-converted voice and the speaker features of the fifth characteristic-converted voice, and the similarity between the speaker features of the fourth characteristic-converted voice and the speaker features of the fifth characteristic-converted voice are calculated.

[0081] The threshold calculation unit 17 calculates a threshold used to recognize the speaker of the input voice, based on the comparison result by the first feature comparison unit 16. More specifically, the threshold calculation unit 17 calculates a threshold based on the multiple similarities calculated by the first feature comparison unit 16. The threshold calculation unit 17 calculates the smallest similarity among the multiple similarities calculated by the first feature comparison unit 16 as the threshold. The threshold calculation unit 17 stores the calculated threshold in the registered voice storage unit 18 in association with the registered voice.

[0082] The registration voice storage unit 18 stores the speaker ID, the registration voice, and the threshold value in association with each other.

[0083] The input speech acquisition unit 19 acquires the input speech of the speaker to be recognized. The input speech acquisition unit 19 may be connected to a microphone (not shown). In the speaker recognition process, the microphone outputs the speech uttered by the speaker to be recognized to the speaker recognition device 1. In the speaker recognition process, the input speech acquisition unit 19 acquires the input speech of the speaker to be recognized from the microphone. The input speech acquisition unit 19 outputs the acquired input speech to the third feature extraction unit 20. Note that the words in the input speech may be different from the words in the registered speech or may be the same as the words in the registered speech.

[0084] In the speaker recognition process, the third feature extracting unit 20 extracts speaker features from the input speech acquired by the input speech acquiring unit 19. The speaker features are, for example, i-vectors.

[0085] In the speaker recognition process, the fourth feature extracting unit 21 extracts speaker features from the enrollment speech acquired by the enrollment speech acquiring unit 12. The speaker features are, for example, i-vectors.

[0086] The second feature comparing unit 22 calculates the similarity between the speaker features extracted from the input speech and the speaker features extracted from the enrolled speech. The method of calculating the similarity by the second feature comparing unit 22 is the same as the method of calculating the similarity by the first feature comparing unit 16.

[0087] The speaker recognition unit 23 determines whether the speaker of the input voice matches the speaker of the registered voice by determining whether the similarity calculated by the second feature comparison unit 22 is greater than a threshold stored in the registered voice storage unit 18. If the similarity calculated by the second feature comparison unit 22 is greater than the threshold stored in the registered voice storage unit 18, the speaker recognition unit 23 determines that the speaker of the input voice matches the speaker of the registered voice. On the other hand, if the similarity calculated by the second feature comparison unit 22 is equal to or less than the threshold stored in the registered voice storage unit 18, the speaker recognition unit 23 determines that the speaker of the input voice does not match the speaker of the registered voice.

[0088] The recognition result output unit 24 outputs a recognition result indicating whether the speaker of the input voice matches the speaker of the registered voice. If the similarity calculated by the second feature comparison unit 22 is greater than a threshold stored in the registered voice storage unit 18, the recognition result output unit 24 outputs a recognition result indicating that the speaker of the input voice matches the speaker of the registered voice. If the similarity calculated by the second feature comparison unit 22 is equal to or less than the threshold stored in the registered voice storage unit 18, the recognition result output unit 24 outputs a recognition result indicating that the speaker of the input voice does not match the speaker of the registered voice.

[0089] The recognition result output unit 24 may output the recognition result to an output device. The output device is, for example, a display or a speaker. If the speaker of the input voice to be recognized is recognized, the output device may output a message indicating that the speaker of the input voice to be recognized is a pre-registered speaker. On the other hand, if the speaker of the input voice to be recognized is not recognized, the output device may output a message indicating that the speaker of the input voice to be recognized is not a pre-registered speaker. Furthermore, the recognition result output unit 24 may output the recognition result by the speaker recognition unit 23 to a device other than the speaker recognition device 1.

[0090] Next, the operation of the voice registration process of the speaker recognition device 1 in this embodiment will be described.

[0091] FIG. 3 is a flowchart for explaining the operation of the speech registration process of the speaker recognition device 1 in this embodiment.

[0092] First, in step S1, the speaker ID acquisition unit 11 acquires a speaker ID for recognizing a speaker to be registered. The speaker ID acquisition unit 11 outputs the acquired speaker ID to the registration voice acquisition unit 12.

[0093] Next, in step S2, the registration speech acquisition unit 12 acquires the registration speech of the speaker to be registered from the microphone. The registration speech acquisition unit 12 outputs the acquired registration speech to the first feature extraction unit 14 and the characteristic conversion unit 13.

[0094] Next, in step S3, the characteristic conversion unit 13 converts the registered voice acquired by the registered voice acquisition unit 12 into a plurality of characteristic-converted voices each having different acoustic characteristics. Here, the first characteristic conversion unit 131 generates first characteristic-converted voice by converting the voice quality of the registered voice into a voice quality changed due to the influence of a cold. Furthermore, the second characteristic conversion unit 132 generates second characteristic-converted voice by converting the speech content of the registered voice into a different speech content. Furthermore, the third characteristic conversion unit 133 generates third characteristic-converted voice by adding noise to the registered voice. Furthermore, the fourth characteristic conversion unit 134 generates fourth characteristic-converted voice by changing the speech rate of the registered voice to a different speech rate. Furthermore, the fifth characteristic conversion unit 135 generates fifth characteristic-converted voice by converting the voice quality of the registered voice into a voice quality expressing a different emotion. The characteristic conversion unit 13 outputs the plurality of characteristic-converted voices (first characteristic-converted voice to fifth characteristic-converted voice) converted from the registered voice to the second feature extraction unit 15.

[0095] Next, in step S4, the first feature extraction unit 14 extracts speaker features from the registration speech acquired by the registration speech acquisition unit 12. The first feature extraction unit 14 outputs the speaker features extracted from the registration speech to the first feature comparison unit 16.

[0096] Next, in step S5, the second feature extraction unit 15 extracts speaker features from each of the multiple characteristic-converted speeches converted by the characteristic conversion unit 13. Here, the fifth feature extraction unit 151 extracts speaker features from the first characteristic-converted speech converted by the first characteristic conversion unit 131. The sixth feature extraction unit 152 extracts speaker features from the second characteristic-converted speech converted by the second characteristic conversion unit 132. The seventh feature extraction unit 153 extracts speaker features from the third characteristic-converted speech converted by the third characteristic conversion unit 133. The eighth feature extraction unit 154 extracts speaker features from the fourth characteristic-converted speech converted by the fourth characteristic conversion unit 134. The ninth feature extraction unit 155 extracts speaker features from the fifth characteristic-converted speech converted by the fifth characteristic conversion unit 135. The second feature extracting unit 15 outputs to the first feature comparing unit 16 a plurality of speaker features extracted from each of the plurality of characteristic-converted speeches (first characteristic-converted speech to fifth characteristic-converted speech).

[0097] Next, in step S6, the first feature comparing unit 16 calculates the similarity of all combinations of two speaker features among the plurality of speaker features extracted from the enrolled voice and the plurality of characteristic-converted voices.

[0098] Next, in step S7, the threshold calculation unit 17 calculates a threshold based on the multiple similarities calculated by the first feature amount comparison unit 16. Here, the threshold calculation unit 17 calculates the smallest similarity among the multiple similarities calculated by the first feature amount comparison unit 16 as the threshold.

[0099] Next, in step S8, the threshold calculation unit 17 stores the calculated threshold in the enrollment voice storage unit 18 in association with the speaker ID and the enrollment voice.

[0100] In the present embodiment, the first feature comparing unit 16 calculates the similarity between all combinations of two speaker features among all of the speaker features extracted from the enrollment speech and the plurality of speaker features extracted from the plurality of characteristic-converted speeches, but the present disclosure is not particularly limited to this. The first feature comparing unit 16 may calculate the similarity between all combinations of two speaker features among some of the speaker features extracted from the enrollment speech and the plurality of speaker features extracted from the plurality of characteristic-converted speeches. The first feature comparing unit 16 may randomly select some of the speaker features extracted from the enrollment speech and the plurality of speaker features extracted from the plurality of characteristic-converted speeches.

[0101] Furthermore, in this embodiment, the first feature comparing unit 16 may calculate the similarity between a speaker feature extracted from the enrollment voice and each of the plurality of speaker features extracted from the plurality of characteristic-converted voices. That is, the first feature comparing unit 16 may calculate the similarity between the speaker feature of the enrollment voice and the speaker feature of the first characteristic-converted voice, the similarity between the speaker feature of the enrollment voice and the speaker feature of the second characteristic-converted voice, the similarity between the speaker feature of the enrollment voice and the speaker feature of the third characteristic-converted voice, the similarity between the speaker feature of the enrollment voice and the speaker feature of the fourth characteristic-converted voice, and the similarity between the speaker feature of the enrollment voice and the speaker feature of the fifth characteristic-converted voice.

[0102] In the present embodiment, the first feature comparing unit 16 may calculate the similarity of all combinations of two speaker features from among all of the plurality of speaker features extracted from the plurality of characteristic-converted speeches. That is, the first feature comparison unit 16 may calculate the similarity between the speaker features of the first characteristic-converted voice and the speaker features of the second characteristic-converted voice, the similarity between the speaker features of the first characteristic-converted voice and the speaker features of the third characteristic-converted voice, the similarity between the speaker features of the first characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, the similarity between the speaker features of the first characteristic-converted voice and the speaker features of the fifth characteristic-converted voice, the similarity between the speaker features of the second characteristic-converted voice and the speaker features of the third characteristic-converted voice, the similarity between the speaker features of the second characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, the similarity between the speaker features of the second characteristic-converted voice and the speaker features of the fifth characteristic-converted voice, the similarity between the speaker features of the third characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, the similarity between the speaker features of the third characteristic-converted voice and the speaker features of the fifth characteristic-converted voice, the similarity between the speaker features of the third characteristic-converted voice and the speaker features of the fourth characteristic-converted voice, and the similarity between the speaker features of the fourth characteristic-converted voice and the speaker features of the fifth characteristic-converted voice.

[0103] In the present embodiment, the first feature comparing unit 16 may calculate the similarity of all combinations of two speaker features among a part of the speaker features extracted from the characteristic-converted speeches.

[0104] In the present embodiment, the threshold calculation unit 17 calculates the minimum similarity among the multiple similarities calculated by the first feature amount comparison unit 16 as the threshold, but the present disclosure is not particularly limited to this. The threshold calculation unit 17 may also calculate the average of the multiple similarities calculated by the first feature amount comparison unit 16 as the threshold.

[0105] Furthermore, the threshold calculation unit 17 may calculate a cumulative distribution function F(x) indicating the proportion of similarity values ​​that are equal to or less than x among the multiple similarities calculated by the first feature comparison unit 16, and may calculate the similarity value when the calculated cumulative distribution function F(x) is a predetermined proportion as the threshold.

[0106] 4 is a diagram showing an example of a cumulative distribution function F(x) calculated based on a plurality of similarities. In FIG. 4, the vertical axis represents the cumulative probability (0 to 100%), and the horizontal axis represents the random variable (similarity value).

[0107] For example, the threshold calculation unit 17 may calculate the similarity value x' when the calculated cumulative distribution function F(x) is 30% as the threshold. Note that the threshold calculation unit 17 may also calculate the similarity value x' when the calculated cumulative distribution function F(x) is any ratio between 20% and 40% as the threshold.

[0108] Next, the operation of the speaker recognition process of the speaker recognition device 1 in this embodiment will be described.

[0109] FIG. 5 is a flowchart for explaining the operation of the speaker recognition process of the speaker recognition device 1 in this embodiment.

[0110] First, in step S11, the speaker ID acquisition unit 11 acquires a speaker ID for recognizing a speaker to be recognized. The speaker ID acquisition unit 11 outputs the acquired speaker ID to the registration voice acquisition unit 12.

[0111] Next, in step S12, the input speech acquisition unit 19 acquires the input speech of the speaker to be recognized from the microphone. The input speech acquisition unit 19 outputs the acquired input speech to the third feature amount extraction unit 20.

[0112] Next, in step S13, the registration speech acquisition unit 12 acquires the registration speech corresponding to the speaker ID acquired by the speaker ID acquisition unit 11 from the registration speech storage unit 18. The registration speech acquisition unit 12 outputs the acquired registration speech to the fourth feature extraction unit 21.

[0113] Next, in step S14, the third feature extraction unit 20 extracts speaker features from the input speech acquired by the input speech acquisition unit 19. The third feature extraction unit 20 outputs the speaker features extracted from the input speech to the second feature comparison unit 22.

[0114] Next, in step S15, the fourth feature extraction unit 21 extracts speaker features from the registration speech acquired by the registration speech acquisition unit 12. The fourth feature extraction unit 21 outputs the speaker features extracted from the registration speech to the second feature comparison unit 22.

[0115] Next, in step S16, the second feature comparison unit 22 calculates the similarity between the speaker features extracted from the input voice by the third feature extraction unit 20 and the speaker features extracted from the registered voice by the fourth feature extraction unit 21.

[0116] Next, in step S17, the speaker recognition unit 23 reads out the threshold value associated with the speaker ID acquired by the speaker ID acquisition unit 11 and the registration voice acquired by the registration voice acquisition unit 12 from the registration voice storage unit 18.

[0117] Next, in step S18, the speaker recognition unit 23 determines whether the similarity calculated by the second feature comparison unit 22 is greater than the threshold value read from the registered voice storage unit 18. If it is determined that the similarity is greater than the threshold value (YES in step S18), in step S19, the recognition result output unit 24 outputs a recognition result indicating that the speaker of the input voice matches the speaker of the registered voice.

[0118] On the other hand, if it is determined that the similarity is below the threshold (NO in step S18), in step S20, the recognition result output unit 24 outputs a recognition result indicating that the speaker of the input voice does not match the speaker of the registered voice.

[0119] In this way, in this embodiment, the registered voice is converted into a plurality of characteristic-converted voices, each having different acoustic characteristics, and all combinations of two speaker features, selected from some or all of the speaker features extracted from the registered voice and the plurality of characteristic-converted voices, are compared, and a threshold value to be used for recognizing the speaker of the input voice is calculated based on the comparison results.

[0120] Therefore, even if the similarity between the input voice and the registered voice decreases due to a change in the acoustic characteristics of the speaker's input voice, it is possible to determine using a threshold whether the decrease in similarity is within an acceptable range depending on the change in acoustic characteristics, thereby reducing the deterioration of speaker recognition performance due to changes in the acoustic characteristics of the speaker's input voice.

[0121] Furthermore, by generating a plurality of characteristic-converted speeches with different acoustic characteristics from the registered speech and calculating the similarity between a plurality of speaker features extracted from the plurality of characteristic-converted speeches, the range of change in the similarity between the speaker features in response to changes in acoustic characteristics such as a change in the speaker's physical condition can be estimated. Therefore, in the speaker recognition process, even if the acoustic characteristics of the input speech change and the similarity between the registered speech and the input speech decreases, it is determined whether the decrease in similarity is due to the change in acoustic characteristics and is acceptable, thereby enabling speaker recognition that is robust against changes in acoustic characteristics.

[0122] In this embodiment, the speaker recognition unit 23 determines whether the speaker of the input voice matches the speaker of the registered voice, but the present disclosure is not limited to this, and the speaker recognition unit 23 may determine which of multiple registered voices the speaker of the input voice is.

[0123] In this case, the registration voice storage unit 18 stores the registration voices of each of the multiple speakers. In the speaker recognition process, the registration voice acquisition unit 12 acquires the multiple registration voices stored in advance from the registration voice storage unit 18. The fourth feature extraction unit 21 extracts speaker features from each of the multiple registration voices acquired by the registration voice acquisition unit 12. The second feature comparison unit 22 calculates similarities between the speaker features extracted from the input voice and each of the multiple speaker features extracted from the multiple registration voices. The speaker recognition unit 23 may determine whether the speaker of the input voice is a pre-registered speaker by determining whether the largest similarity among the multiple similarities calculated by the second feature comparison unit 22 is greater than a threshold stored in the registration voice storage unit 18.

[0124] That is, when it is determined that the largest similarity among the multiple similarities calculated by the second feature comparing unit 22 is greater than the threshold value stored in the registered voice storage unit 18, the speaker recognizing unit 23 determines that the speaker of the input voice is a pre-registered speaker. On the other hand, when it is determined that the largest similarity among the multiple similarities calculated by the second feature comparing unit 22 is equal to or smaller than the threshold value stored in the registered voice storage unit 18, the speaker recognizing unit 23 determines that the speaker of the input voice is not a pre-registered speaker. Then, the recognition result output unit 24 may output a recognition result indicating whether the speaker of the input voice matches any of the speakers of the multiple registered voices.

[0125] Furthermore, in this embodiment, when the speaker recognition device 1 recognizes only one pre-registered speaker, a speaker ID is not necessary, and the speaker recognition device 1 does not need to include the speaker ID acquisition unit 11.

[0126] In each of the above embodiments, each component may be configured with dedicated hardware or may be realized by executing a software program suitable for that component. Each component may be realized by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording the program on a recording medium and transferring it, or by transferring the program via a network.

[0127] Some or all of the functions of the device according to the embodiments of the present disclosure are typically realized as an LSI (Large Scale Integration), which is an integrated circuit. These may be implemented individually on a single chip, or some or all of them may be integrated on a single chip. Furthermore, the integrated circuit is not limited to an LSI, and may be realized using a dedicated circuit or a general-purpose processor. It is also possible to use an FPGA (Field Programmable Gate Array), which can be programmed after LSI manufacturing, or a reconfigurable processor, which allows the connections and settings of circuit cells within the LSI to be reconfigured.

[0128] Furthermore, some or all of the functions of the device according to the embodiment of the present disclosure may be realized by a processor such as a CPU executing a program.

[0129] Furthermore, all the numbers used above are merely examples to specifically explain the present disclosure, and the present disclosure is not limited to the numbers used as examples.

[0130] The order in which the steps are performed shown in the above flowchart is merely an example for specifically explaining the present disclosure, and other orders may be used as long as similar effects are obtained. Also, some of the steps may be performed simultaneously (in parallel) with other steps. [Industrial Applicability]

[0131] The technology disclosed herein can reduce degradation of speaker recognition performance due to changes in the acoustic characteristics of a speaker's input voice, and is therefore useful as a technology for recognizing the speaker of an input voice by comparing the input voice with a registered voice.

Claims

1. The computer Get the registration voice, converting the acquired registered voice into a plurality of characteristic-converted voices each having different acoustic characteristics; extracting speaker features indicating speaker characteristics from the registered voice; extracting the speaker features from each of the plurality of characteristic-converted voices; comparing all combinations of two speaker features among a part or all of the speaker features extracted from the enrollment speech and the speaker features extracted from the characteristic-converted speeches; calculating a threshold value to be used for recognizing the speaker of the input speech based on the comparison result; Information processing methods.

2. Furthermore, the threshold value is stored in a memory in association with the registered voice; Furthermore, the input speech of the speaker to be recognized is acquired, Furthermore, the registration voice is acquired, Furthermore, the speaker features are extracted from the input speech; Furthermore, the speaker features are extracted from the enrollment voices, Furthermore, a similarity between the speaker feature extracted from the input speech and the speaker feature extracted from the enrollment speech is calculated; and outputting a recognition result indicating that the speaker of the input voice matches the speaker of the registered voice when the calculated similarity is greater than the threshold stored in the memory. The information processing method according to claim 1.

3. In comparing all combinations of the two speaker features, a similarity is calculated for all combinations of the two speaker features, among the speaker features extracted from the enrollment speech and some or all of the plurality of speaker features extracted from the plurality of characteristic-converted speeches; In calculating the threshold value, the threshold value is calculated based on the calculated multiple similarities.

3. The information processing method according to claim 1 or 2.

4. In calculating the threshold, the minimum similarity among the calculated similarities is calculated as the threshold.

4. The information processing method according to claim 3.

5. In calculating the threshold, an average of the calculated similarities is calculated as the threshold.

4. The information processing method according to claim 3.

6. In calculating the threshold, a cumulative distribution function F(x) indicating a proportion of the calculated similarities whose similarity values ​​are equal to or less than x is calculated, and the value of the similarity when the calculated cumulative distribution function F(x) is a predetermined proportion is calculated as the threshold.

4. The information processing method according to claim 3.

7. the plurality of characteristic-converted voices include a first characteristic-converted voice obtained by converting the voice quality of the registered voice into a voice quality changed due to the influence of a cold; 3. The information processing method according to claim 1 or 2.

8. the plurality of characteristic-converted voices include a second characteristic-converted voice obtained by converting the speech content of the registration voice into a different speech content; 3. The information processing method according to claim 1 or 2.

9. the plurality of characteristic-converted sounds include a third characteristic-converted sound obtained by adding noise to the registered sound; 3. The information processing method according to claim 1 or 2.

10. the plurality of characteristic-converted voices include a fourth characteristic-converted voice obtained by converting the speech rate of the registered voice to a different speech rate; 3. The information processing method according to claim 1 or 2.

11. the plurality of characteristic-converted voices include a fifth characteristic-converted voice obtained by converting the voice quality of the registered voice into a voice quality that expresses a predetermined emotion; 3. The information processing method according to claim 1 or 2.

12. an acquisition unit for acquiring a registration voice; a conversion unit that converts the acquired registered voice into a plurality of characteristic-converted voices each having a different acoustic characteristic; a first extraction unit that extracts speaker features indicating speaker characteristics from the registered voice; a second extraction unit that extracts the speaker features from each of the plurality of characteristic-converted speeches; a comparison unit that compares all combinations of two speaker features among a speaker feature extracted from the enrollment speech and some or all of the speaker features extracted from the characteristic-converted speeches; a calculation unit that calculates a threshold value used to recognize the speaker of the input voice based on the comparison result; An information processing device comprising:

13. Get the registration voice, converting the acquired registered voice into a plurality of characteristic-converted voices each having different acoustic characteristics; extracting speaker features indicating speaker characteristics from the registered voice; extracting the speaker features from each of the plurality of characteristic-converted voices; comparing all combinations of two speaker features among a part or all of the speaker features extracted from the enrollment speech and the speaker features extracted from the characteristic-converted speeches; An information processing program that causes a computer to calculate a threshold value used to recognize the speaker of the input voice based on the comparison result.

Citation Information

Patent Citations

  • Voice recognition equipment

    JP1987124599A

  • Speaker collating device, its method and storage medium

    JP1999327586A

  • Method, apparatus, and system for speaker verification

    JP2019527370A

  • Training method, speaker identification method, and recording medium

    JP2021033260A

  • Speaker recognizing device, speaker recognizing method, etc.

    WO2008018136A1