Learning method, speaker recognition method, and recording medium

By using a learning method for a speaker recognition model, the voice data of a first speaker is transformed to generate the voice data of a second speaker, thereby solving the problem of the influence of speech content and language in the prior art and achieving high-precision speaker recognition.

CN112420021BActive Publication Date: 2025-09-19PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010829027.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-24
Filing Date
2020-08-18
Publication Date
2025-09-19
Estimated Expiration
2040-08-18

AI Technical Summary

Technical Problem

The technical problem in the existing technology of speaker recognition model is how to improve speaker recognition. The existing technology cannot effectively identify the amount of learning data of the model and cannot fully reduce the influence of the voice content and language.

Method used

By transforming the voice characteristics of the first speaker to generate the voice data of the second speaker, the learning data is processed by adding noise, adding reverberation, etc., and the voice data of the second speaker is generated to increase the amount of learning data.

Benefits of technology

The invention realizes increasing the amount of the second speaker's voice data in the learning process of the speaker recognition model without being restricted by the utterance content and language, thereby improving the accuracy of speaker recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112420021B_ABST
    Figure CN112420021B_ABST
Patent Text Reader

Abstract

The problem to be solved by the present invention is to identify a speaker with high accuracy. A learning method, a speaker identification method, and a recording medium are provided. The learning method is a learning method for a speaker identification model (20). When voice data is input, the speaker identification model (20) outputs speaker identification information identifying the speaker uttering the voice contained in the voice data. The speaker identification model (20) generates second voice data of a second speaker by performing voice feature transformation processing on first voice data of a first speaker. The first voice data and the second voice data are used as learning data for learning processing of the speaker identification model (20).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to techniques for identifying speakers. Background Art

[0002] Conventionally, there is known a technique for identifying a speaker using a speaker identification model (for example, see Non-Patent Document 1).

[0003] Prior art literature

[0004] Non-patent literature

[0005] Non-patent literature 1: David Snyder, Daniel Garcia-Romero, Gregory Sell, Daniel Povey, Sanjeev Khudanpur, "X-VECTORS: ROBUST DNN EMBEDDINGS FOR SPEAKERRECOGNITION" ICASSP 2018: 5329-5333. Summary of the Invention

[0006] Problems to be solved by the invention

[0007] It is desirable to identify the speaker with high accuracy.

[0008] Means for solving problems

[0009] One aspect of the present disclosure is a learning method for a speaker recognition model, wherein when sound data is input, the speaker recognition model outputs speaker recognition information for identifying the speaker uttering the sound contained in the sound data, wherein second sound data of a second speaker is generated by performing sound feature transformation processing on first sound data of a first speaker, and the first sound data and the second sound data are used as learning data for learning processing of the speaker recognition model.

[0010] A speaker identification method according to one aspect of the present disclosure inputs voice data to the speaker identification model that has been previously trained by the above-mentioned learning method, and causes the speaker identification model to output the speaker identification information.

[0011] A recording medium according to one embodiment of the present disclosure is a computer-readable recording medium having a program recorded thereon, wherein the program is used to cause a computer to execute processing for learning a speaker recognition model, wherein the speaker recognition model, when input with sound data, outputs speaker recognition information identifying the speaker uttering the sound contained in the sound data, wherein the processing includes: a first step of generating second sound data of a second speaker by performing sound feature transformation processing on first sound data of a first speaker; and a second step of performing learning processing on the speaker recognition model using the first sound data and the second sound data as learning data.

[0012] In addition, these overall or specific methods can be implemented through systems, methods, integrated circuits, computer programs or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs and recording media.

[0013] Effects of the Invention

[0014] According to the learning method and the like of the present disclosure, it is possible to identify a speaker with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a block diagram showing a configuration example of a speaker identification device according to an embodiment.

[0016] Figure 2 This is a schematic diagram showing an example of a situation in which the audio data storage unit according to the embodiment stores audio data and speaker identification information in association with each other.

[0017] Figure 3 This is a schematic diagram showing a situation in which the voice characteristic conversion unit of the embodiment converts voice data of one speaker into voice data of a plurality of other speakers and outputs the converted data.

[0018] Figure 4 This is a block diagram showing a configuration example of a voice characteristic conversion unit according to an embodiment.

[0019] Figure 5 This is a flowchart of the speaker recognition model learning process according to the embodiment.

[0020] Figure 6 This is a flowchart of the voice characteristic conversion model learning process of the implementation method.

[0021] Figure 7 This is a flowchart of speaker recognition processing according to an embodiment.

[0022] Description of Reference Numerals

[0023] 1 Speaker Recognition Device

[0024] 10 Sound data extension unit

[0025] 11. Voice data storage unit

[0026] 12. First audio data acquisition unit

[0027] 13. Voice Characteristic Transformation Unit

[0028] 14 Noise echo imparting unit

[0029] 15. First feature quantity calculation unit

[0030] 16 Comparative Department

[0031] 17 Voice data storage unit

[0032] 18 Extended audio data storage unit

[0033] 20 Speaker Identification Model

[0034] 21 Third feature quantity calculation unit

[0035] 22 Deep Neural Networks

[0036] 23 Judgment Department

[0037] 30 Learning Department

[0038] 31 Second audio data acquisition unit

[0039] 32 Second feature quantity calculation unit

[0040] 33 First Learning Department

[0041] 40 Identification target sound data acquisition unit

[0042] 131 Voice characteristic transformation learning data storage unit

[0043] 132 Second Learning Department

[0044] 133 Sound Transformation Model DETAILED DESCRIPTION

[0045] (Process of Arriving at One Mode of the Present Disclosure)

[0046] Speaker recognition technology is known that identifies a speaker using a speaker recognition model that has been previously trained using, as learning data, voice data associated with recognition information for identifying a speaker.

[0047] In the past, to increase the amount of learning data (hereinafter, "increasing the amount of learning data" is also referred to as "expanding the learning data"), the original learning sound data was subjected to noise addition, reverberation, and other methods. However, with these conventional methods of expanding learning data, it is not possible to increase the content of a speaker's speech or the language (Japanese, English, etc.). Consequently, there is a situation where the influence of speech content and language on the learning process of the speaker recognition model cannot be fully reduced.

[0048] Therefore, the inventors have conducted extensive research and experiments to achieve high-precision speaker recognition using a speaker recognition model, and have come up with the following learning method and the like.

[0049] One aspect of the present disclosure is a learning method for a speaker recognition model, wherein when sound data is input, the speaker recognition model outputs speaker recognition information for identifying the speaker uttering the sound contained in the sound data, wherein second sound data of a second speaker is generated by performing sound feature transformation processing on first sound data of a first speaker, and the first sound data and the second sound data are used as learning data for learning processing of the speaker recognition model.

[0050] According to the above learning method, when expanding the learning data in the speaker identification model learning process, the amount of voice data of the second speaker can be increased without being restricted by the utterance content or language. Therefore, the speaker identification accuracy of the speaker identification model can be improved.

[0051] Therefore, according to the above-mentioned learning method, the speaker can be identified with high accuracy.

[0052] Furthermore, the voice characteristic conversion process may be a process based on the voice data of the first speaker and the voice data of the second speaker.

[0053] Alternatively, the voice characteristic transformation processing may include inputting the first voice data into a voice characteristic transformation model, thereby outputting the second voice data from the voice characteristic transformation model, and the voice characteristic transformation model may be pre-learned so that when the voice data of the first speaker is input, the voice data of the second speaker is output.

[0054] Alternatively, the sound characteristic transformation model may include a deep neural network that takes sound data in WAV format as input and outputs sound data in WAV format.

[0055] Furthermore, the voice characteristic conversion process may be a process based on voice data of the first speaker and voice data of a third speaker.

[0056] Furthermore, the speaker recognition model may include a deep neural network that receives as input an utterance feature quantity representing a feature of an utterance included in the audio data and outputs a speaker characteristic feature quantity representing a feature of a speaker.

[0057] In a speaker identification method according to one aspect of the present disclosure, voice data is input to the speaker identification model that has been previously trained by the above-mentioned learning method, and the speaker identification model is caused to output the speaker identification information.

[0058] According to the speaker identification method, the amount of second speaker's voice data can be increased during the expansion of learning data in the speaker identification model learning process, regardless of the utterance content or language. This improves the speaker identification accuracy of the speaker identification model.

[0059] Therefore, according to the above-described speaker identification method, it is possible to identify the speaker with high accuracy.

[0060] A recording medium according to one embodiment of the present disclosure is a computer-readable recording medium having a program recorded thereon, wherein the program is used to cause a computer to execute processing for learning a speaker recognition model, wherein the speaker recognition model, when input with sound data, outputs speaker recognition information identifying the speaker uttering the sound contained in the sound data, wherein the processing includes: a first step of generating second sound data of a second speaker by performing sound feature transformation processing on first sound data of a first speaker; and a second step of performing learning processing on the speaker recognition model using the first sound data and the second sound data as learning data.

[0061] According to the recording medium, the amount of second speaker's voice data can be increased during the expansion of learning data in the speaker identification model learning process, regardless of the utterance content or language. Therefore, the speaker identification accuracy of the speaker identification model can be improved.

[0062] Therefore, according to the above-mentioned recording medium, it is possible to identify a speaker with high accuracy.

[0063] In addition, these general or specific methods can be implemented through systems, methods, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0064] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. Each embodiment described below represents a specific example of the present disclosure. The numerical values, shapes, constituent elements, steps, and the order of steps shown in the following embodiments are examples and are not intended to limit the present disclosure. In addition, in all embodiments, the respective contents may also be combined.

[0065] (Implementation Method)

[0066] Hereinafter, a speaker identification device according to an embodiment will be described. This speaker identification device acquires voice data and outputs identification information identifying a speaker who uttered a voice contained in the voice data.

[0067] <Structure>

[0068] Figure 1 This is a block diagram showing a configuration example of the speaker identification device 1 according to the embodiment.

[0069] like Figure 1 As shown, the speaker identification apparatus 1 includes a speech data expansion unit 10 , a speaker identification model 20 , a learning unit 30 , and a recognition target speech data acquisition unit 40 .

[0070] The audio data expansion unit 10 expands the learning data used to learn the speaker recognition model 20 (i.e., increases the amount of learning data). The audio data expansion unit 10 can also be implemented, for example, by a computer equipped with a microprocessor, memory, a communication interface, etc. In this case, the various functions of the audio data expansion unit 10 are implemented by the microprocessor executing a program stored in the memory. Alternatively, the audio data expansion unit 10 can be implemented, for example, by distributed computing or cloud computing performed by multiple computers communicating with each other.

[0071] like Figure 1 As shown, the audio data expansion unit 10 includes an audio data storage unit 11, a first audio data acquisition unit 12, an audio characteristic conversion unit 13, a noise reverberation imparting unit 14, a first feature quantity calculation unit 15, a comparison unit 16, an audio data storage unit 17, and an expanded audio data storage unit 18.

[0072] The learning unit 30 uses the learning data expanded by the speech data expansion unit 10 to perform a learning process for the speaker recognition model 20. The learning unit 30 can also be implemented, for example, by a computer equipped with a microprocessor, memory, a communication interface, and the like. In this case, the various functions of the learning unit 30 are implemented by the microprocessor executing a program stored in the memory. Alternatively, the learning unit 30 can be implemented, for example, by distributed computing or cloud computing performed by multiple computers communicating with each other.

[0073] like Figure 1As shown, the learning unit 30 includes a second audio data acquisition unit 31 , a second feature quantity calculation unit 32 , and a first learning unit 33 .

[0074] When audio data is input, speaker identification model 20 outputs speaker identification information identifying the speaker uttering the audio data. Speaker identification model 20 can also be implemented, for example, by a computer equipped with a microprocessor, memory, and a communication interface. In this case, the various functions of speaker identification model 20 are implemented by the microprocessor executing a program stored in the memory. Alternatively, speaker identification model 20 can be implemented, for example, by distributed computing or cloud computing performed by multiple computers communicating with each other.

[0075] like Figure 1 As shown, the speaker recognition model 20 includes a third feature quantity calculation unit 21 , a deep neural network (DNN) 22 , and a determination unit 23 .

[0076] The recognition target audio data acquisition unit 40 acquires audio data to be recognized in speaker recognition performed by the speaker recognition model 20. The recognition target audio data acquisition unit 40 may, for example, include a communication interface for communicating with an external device and acquire audio data from the external device via the communication interface. Alternatively, the recognition target audio data acquisition unit 40 may include an input / output port (e.g., a USB port) and acquire audio data from an external storage device (e.g., a USB memory) connected to the input / output port. Alternatively, the recognition target audio data acquisition unit 40 may include a microphone and acquire audio data by converting audio input to the microphone into an electrical signal.

[0077] Hereinafter, each component constituting the audio data expansion unit 10 will be described.

[0078] The audio data storage unit 11 stores audio data and speaker identification information associated with the audio data and identifying a speaker who uttered the audio contained in the audio data, in association with each other.

[0079] Figure 2 This is a schematic diagram showing an example of how the voice data storage unit 11 stores voice data and speaker recognition information in association with each other.

[0080] like Figure 2 As shown, the speech data storage unit 11 stores a plurality of speech data associated with a plurality of different speaker identification information. The speech data and speaker identification information stored in the speech data storage unit 11 are used as learning data for learning the speaker identification model 20 .

[0081] Return again Figure 1 , the description of the speaker recognition apparatus 1 is continued.

[0082] The voice data storage unit 11 may include, for example, a communication interface for communicating with an external device, and store voice data acquired from the external device via the communication interface, and speaker identification information associated with the voice data. Furthermore, the voice data storage unit 11 may include, for example, an input / output port (e.g., a USB port), and store voice data acquired from an external storage device (e.g., a USB memory) connected to the input / output port, and speaker identification information associated with the voice data.

[0083] Here, the audio data is described as being in the WAV format, but the audio data is not necessarily limited to the WAV format, and may be in the AIFF format, AAC format, or the like, for example.

[0084] The first audio data acquisition unit 12 acquires audio data and speaker recognition information associated with the audio data from the audio data storage unit 11 .

[0085] The voice characteristic conversion unit 13 converts the voice data acquired by the first voice data acquisition unit 12 into voice data uttered by a speaker other than the speaker identified by the speaker identification information associated with the voice data (hereinafter also referred to as "the other speaker"), and outputs the converted voice data. More specifically, the voice characteristic conversion unit 13 generates and outputs voice data uttered by the other speaker by modifying the frequency components of the utterance included in the voice data.

[0086] By converting the voice data of a single speaker into the voice data of multiple other speakers and outputting the converted data, the voice characteristic conversion unit 13 can output multiple voice data sets with the same utterance content, even though the speakers are different. Furthermore, if the voice data of a single speaker contains Japanese pronunciations, the voice characteristic conversion unit 13 can convert the data into voice data containing Japanese pronunciations of other speakers who may not necessarily speak Japanese. In other words, the voice characteristic conversion unit 13 can convert the voice data of a single speaker into the voice data of multiple other speakers and output them, regardless of the utterance content or language of the pre-converted voice data.

[0087] Figure 3 1 is a schematic diagram showing a state in which the voice characteristic conversion unit 13 converts voice data of one speaker into voice data of a plurality of other speakers and outputs the converted data.

[0088] like Figure 3 As shown, the voice characteristic conversion unit 13 can increase the amount of voice data used as learning data for the learning process of the speaker recognition model 20 without being restricted by the utterance content or language.

[0089] Return again Figure 1 , the description of the speaker recognition apparatus 1 is continued.

[0090] The voice quality conversion unit 13 can be implemented, for example, by a widely available existing voice quality converter. Alternatively, for example, the voice quality conversion unit 13 can be implemented by utilizing a voice quality conversion model that has been pre-learned so that when voice data of a first speaker is input, voice data of a second speaker is output. Here, the voice quality conversion unit 13 is implemented by utilizing a voice quality conversion model that has been pre-learned so that when voice data of a first speaker is input, voice data of a second speaker is output.

[0091] Figure 4 1 is a block diagram showing a configuration example of the voice characteristic conversion unit 13 .

[0092] like Figure 4 As shown, the voice characteristic conversion unit 13 includes a voice characteristic conversion learning data storage unit 131 , a second learning unit 132 , and a voice characteristic conversion model 133 .

[0093] The voice trait conversion model 133 is a deep neural network (DNN) pre-trained for multiple speaker pairs. When the voice data of a first speaker, one speaker in a speaker pair, is input, the model outputs the voice data of a second speaker, the other speaker in the speaker pair. When the voice data of the second speaker is input, the model outputs the voice data of the first speaker. Here, as an example, the voice trait conversion model 133 is a cycleVAE. The cycleVAE is pre-trained for multiple speaker pairs. When the WAV-formatted voice data of the first speaker is input, the model outputs the WAV-formatted voice data of the second speaker, and when the WAV-formatted voice data of the second speaker is input, the model outputs the WAV-formatted voice data of the first speaker. However, the voice trait conversion model 133 is not necessarily limited to the cycleVAE described above, as long as the DNN is pre-trained for multiple speaker pairs. When the voice data of the first speaker is input, the model outputs the voice data of the second speaker, and when the voice data of the second speaker is input, the model outputs the voice data of the first speaker.

[0094] The voice characteristic conversion learning data storage unit 131 stores learning data used for learning the voice characteristic conversion model 133. More specifically, the voice characteristic conversion learning data storage unit 131 stores voice data (here, voice data in WAV format) of each of a plurality of speakers for which the voice characteristic conversion model 133 is a target.

[0095] The second learning unit 132 uses the learning data stored in the voice feature transformation learning data holding unit 131 to perform learning processing of the voice feature transformation model 133 on multiple speaker pairs, so that when the voice data of the first speaker who is one speaker of the speaker pair is input, the voice data of the second speaker who is the other speaker of the speaker pair is output, and when the voice data of the second speaker is input, the voice data of the first speaker is output.

[0096] Return again Figure 1 , the description of the speaker recognition apparatus 1 is continued.

[0097] The noise and reverberation imparting unit 14 imparts noise (e.g., four types) and reverberation (e.g., one type) to each of the sound data output from the sound characteristic conversion unit 13, and outputs the noise-imparted sound data and the reverberation-imparted sound data. This allows the noise and reverberation imparting unit 14 to further increase the amount of sound data.

[0098] The first feature quantity calculation unit 15 calculates, based on the sound data output from the sound characteristic conversion unit 13 and the sound data output from the noise reverberation imparting unit 14, speech feature quantities representing the characteristics of the speech contained in the sound data. Here, as an example, the first feature quantity calculation unit 15 calculates MFCCs (Mel-Freuyency Cepstrum Coefficients) representing the characteristics of the speaker's vocal tract as speech feature quantities. However, the first feature quantity calculation unit 15 is not necessarily limited to calculating MFCCs, as long as it can calculate speech feature quantities representing the characteristics of the speaker. For example, the first feature quantity calculation unit 15 may calculate the speech feature quantity by applying a Mel filter bank to the speech signal, or may calculate the speech feature quantity by using the spectrum of the speech signal.

[0099] The comparison unit 16 compares the speaker feature value output from the first feature value calculation unit 15 (hereinafter also referred to as the "first speaker feature value") with the speaker feature value of the speaker who uttered the speech included in the sound data that serves as the calculation source of the first speaker feature value (hereinafter also referred to as the "second speaker feature value").

[0100] When the comparison result shows that (1) the similarity between the first speaker feature and the second speaker feature is within a predetermined range, the comparison unit 16 associates the voice data used as the calculation source of the first speaker feature with the speaker identification information identifying the speaker of the utterance included in the voice data. Thus, the comparison unit 16 can increase the amount of voice data associated with a single piece of speaker identification information. The comparison unit 16 then outputs the voice data and the speaker identification information associated with the voice data.

[0101] When the comparison result is (2) that the similarity between the first speaker feature quantity and the second speaker feature quantity is not within a prescribed range, the comparison unit 16 associates the voice data that serves as the calculation source of the first speaker feature quantity with identification information for identifying a third person different from the speaker uttered by the voice data. Thus, the comparison unit 16 can increase the number of speaker identification information associated with the voice data. That is, the comparison unit 16 can increase the number of speakers in the learning data used for the learning process of the speaker recognition model 20. By increasing the number of speakers, over-learning in the learning process of the speaker recognition model 20 described later can be suppressed. Thus, the generalization performance of the speaker recognition model 20 can be improved. Then, the comparison unit 16 outputs the voice data and the speaker identification information associated with the voice data.

[0102] Similar to the voice data storage unit 11 , the extended voice data storage unit 18 stores voice data and speaker identification information associated with the voice data and identifying the speaker of the utterance included in the voice data, in association with each other.

[0103] The audio data storage unit 17 stores the audio data output from the comparison unit 16 and the speaker identification information associated with the audio data in the extended audio data storage unit 18, respectively, in association with each other. Furthermore, the audio data storage unit 17 stores the audio data acquired by the first audio data acquisition unit 12 and the speaker identification information associated with the audio data in the extended audio data storage unit 18, respectively, in association with each other. Thus, in addition to the audio data stored by the audio data storage unit 11 as learning data for learning the speaker identification model 20, the extended audio data storage unit 18 also stores the audio data output from the comparison unit 16 as learning data for learning the speaker identification model.

[0104] Hereinafter, each component constituting the speaker identification model 20 will be described.

[0105] The third feature quantity calculation unit 21 calculates, based on the sound data acquired by the recognition target sound data acquisition unit 40, a vocalization feature quantity representing the characteristics of the vocalization contained in the sound data. Here, as an example, the third feature quantity calculation unit 21 calculates MFCCs representing the vocal tract characteristics of the speaker as the vocalization feature quantity. However, the third feature quantity calculation unit 21 is not necessarily limited to calculating MFCCs, as long as it can calculate vocalization feature quantities representing the characteristics of the speaker. For example, the third feature quantity calculation unit 21 may calculate the value obtained by applying a Mel filter bank to the vocalized sound signal as the vocalization feature quantity, or, for example, may calculate the spectrum of the vocalized sound signal as the vocalization feature quantity.

[0106] The deep neural network 22 is a deep neural network (DNN) that has been pre-learned so that, when the utterance feature calculated by the third feature calculation unit 21 is input, it outputs speaker-specific features representing the characteristics of the speaker of the utterance included in the sound data that served as the calculation source for the utterance feature. Here, as an example, the deep neural network 22 is described as a Kaldi that has been pre-learned so that, when MFCCs representing the vocal tract characteristics of the speaker are input, an x-Vector, which is an acoustic feature of the utterance mapped from a variable-length utterance to a fixed-dimensional embedding, is output as the speaker-specific feature. However, as long as the deep neural network 22 is a DNN that has been pre-learned so that, when the utterance feature calculated by the third feature calculation unit 21 is input, it outputs speaker-specific features representing the characteristics of the speaker, it is not necessarily limited to the above-mentioned Kaldi. Furthermore, details of the x-Vector calculation method, etc., are disclosed in Non-Patent Document 1, and therefore a detailed description is omitted here.

[0107] Based on the speaker-specific feature output from the deep neural network 22, the determination unit 23 determines the speaker of the speech contained in the speech data obtained by the recognition target speech data acquisition unit 40. More specifically, the determination unit 23 stores x-vectors for multiple speakers, identifies the x-vector most similar to the x-vector output from the deep neural network 22 among the stored x-vectors, and determines the speaker of the identified x-vector as the speaker of the speech contained in the speech data obtained by the recognition target speech data acquisition unit 40. The determination unit 23 then outputs speaker identification information indicating the speaker identified.

[0108] Hereinafter, each component constituting the learning unit 30 will be described.

[0109] The second audio data acquisition unit 31 acquires audio data and speaker recognition information associated with the audio data from the extended audio data storage unit 18 .

[0110] The second feature quantity calculation unit 32 calculates, based on the sound data acquired by the second sound data acquisition unit 31, a vocalization feature quantity representing the characteristics of the vocalization contained in the sound data. Here, as an example, the second feature quantity calculation unit 32 calculates MFCCs representing the characteristics of the speaker's vocal tract as the vocalization feature quantity. However, the second feature quantity calculation unit 32 is not necessarily limited to calculating MFCCs, as long as it can calculate vocalization feature quantities representing the characteristics of the speaker. For example, the second feature quantity calculation unit 32 may calculate the value obtained by applying a Mel filter bank to the vocalized sound signal as the vocalization feature quantity, or, for example, may calculate the spectrum of the vocalized sound signal as the vocalization feature quantity.

[0111] The first learning unit 33 uses the speech feature calculated by the second feature calculation unit 32 and the speaker identification information of the speaker of the speech contained in the sound data that serves as the calculation source of the speech feature as learning data, and performs learning processing on the speaker identification model 20, so that when the sound data is input, the speaker identification information of the speaker of the speech contained in the sound data is output.

[0112] More specifically, the first learning unit 33 uses the MFCC calculated by the second feature value calculation unit 32 and the speaker identification information corresponding to the MFCC as learning data. When the MFCC is input, the deep neural network 22 performs learning processing to output an x-Vector representing the characteristics of the speaker who speaks contained in the sound data that serves as the calculation source of the MFCC.

[0113] <Action>

[0114] The speaker identification apparatus 1 having the above-described structure performs speaker identification model learning processing, voice characteristic conversion model learning processing, and speaker identification processing.

[0115] Hereinafter, these processes will be described in sequence with reference to the drawings.

[0116] Figure 5 It is a flowchart of the speaker recognition model learning process.

[0117] The speaker identification model learning process is a process of learning the speaker identification model 20 .

[0118] The speaker recognition model learning process is started, for example, when a user of the speaker recognition apparatus 1 performs an operation on the speaker recognition apparatus 1 to start the speaker recognition model learning process.

[0119] When the speaker recognition model learning process starts, the first voice data acquisition unit 12 acquires one voice data and one speaker recognition information associated with the voice data from the voice data storage unit 11 (step S100 ).

[0120] When acquiring one piece of audio data and one piece of speaker identification information, the audio data storage unit 17 stores the one piece of audio data and the one piece of speaker identification information in association with each other in the extended audio data storage unit 18 (step S110 ).

[0121] On the other hand, the voice characteristic conversion unit 13 selects a speaker from among the other speakers other than the speaker identified by the speaker identification information (step S120). The voice characteristic conversion unit 13 then converts the one voice data into voice data uttered by the one speaker (step S130) and outputs the data.

[0122] When the voice data is output from the voice characteristic conversion unit 13 , the noise and reverberation imparting unit 14 imparts noise and reverberation to the voice data output from the voice characteristic conversion unit 13 (step S140 ), and outputs one or more voice data.

[0123] When one or more speech data are output from the noise reverberation imparting unit 14 , the first feature quantity calculating unit 15 calculates speech feature quantities based on the speech data output from the speech characteristic converting unit 13 and the one or more speech data output from the noise reverberation imparting unit 14 (step S150 ).

[0124] When the utterance feature is calculated, the comparison unit 16 compares the calculated utterance feature with the utterance feature of the selected speaker and determines whether the similarity between the calculated utterance feature and the utterance feature of the speaker is within a predetermined range (step S160).

[0125] If the comparison unit 16 makes a positive determination in step S160 (step S160: Yes), the speaker identification information identifying the selected speaker is associated with the voice data that served as the source of the utterance feature value for which the positive determination was made (step S170). The comparison unit 16 then outputs the voice data and the associated speaker identification information.

[0126] If the comparison unit 16 makes a negative determination in step S160 (step S160: No), it associates identification information identifying a third person different from the selected speaker with the speech data that served as the source of the utterance feature value for which the negative determination was made (step S180). The comparison unit 16 then outputs the speech data and the speaker identification information associated with the speech data.

[0127] When the comparison unit 16 performs the processing of step S170 or the processing of step S180 for all the speech feature quantities that become the comparison objects in the processing of step S160, the sound data storage unit 17 stores the sound data output from the comparison unit 16 and the speaker identification information associated with the sound data in the extended sound data holding unit 18 in correspondence with each other (step S190).

[0128] When the process of step S190 is completed, the voice characteristic conversion unit 13 determines whether there is a speaker not selected in the process of step S120 (hereinafter also referred to as “unselected speaker”) among the other speakers (step S200 ).

[0129] In the process of step S200 , when it is determined that there is an unselected speaker (step S200 : Yes), the voice characteristic conversion unit 13 selects one speaker from the unselected speakers (step S210 ), and proceeds to the process of step S130 .

[0130] In the process of step S200 , if it is determined that there is no unselected speaker (step S200 : No), the first voice data acquisition unit 12 determines whether there is any unacquired voice data in the voice data stored in the voice data storage unit 11 (step S220 ).

[0131] In the process of step S220 , when it is determined that unacquired audio data exists (step S220 : Yes), the first audio data acquisition unit 12 acquires one audio data from the unacquired audio data (step S230 ), and the process proceeds to the process of step S110 .

[0132] In the process of step S220, when it is determined that there is no unacquired voice data (step S220: No), the second voice data acquisition unit 31 acquires the voice data and the speaker recognition information associated with the voice data from the extended voice data storage unit 18 for all the voice data stored in the extended voice data storage unit 18 (step S240).

[0133] After acquiring the voice data and the speaker identification information associated with the voice data for all voice data, the second feature quantity calculation unit 32 calculates, based on the voice data, utterance feature quantities representing features of utterances included in the voice data (step S250 ).

[0134] If the vocal feature quantities are calculated for all the sound data, the first learning unit 33 performs learning processing of the speaker recognition model 20 on all the vocal feature quantities, using the vocal feature quantities and the speaker identification information of the speaker who speaks included in the sound data that serves as the calculation source of the vocal feature quantities as learning data, so that when the sound data is input, the speaker identification information of the speaker who speaks included in the sound data is output (step S260).

[0135] When the process of step S260 ends, the speaker identification apparatus 1 ends the speaker identification model learning process.

[0136] Figure 6 This is a flowchart of the sound feature transformation model learning process.

[0137] The voice characteristic conversion model learning process is a process of performing a learning process of the voice characteristic conversion model 133 .

[0138] The voice characteristic conversion model learning process is started, for example, when a user of the speaker recognition apparatus 1 performs an operation on the speaker recognition apparatus 1 to start the voice characteristic conversion model learning process.

[0139] When the voice characteristic transformation model learning process begins, the second learning unit 132 selects one of the multiple speakers for which the voice characteristic transformation model 133 is to be trained (step S300). The second learning unit 132 then uses the learning data of each of the two speakers constituting the selected speaker pair, stored in the voice characteristic transformation learning data storage unit 131, to perform voice characteristic transformation model 133 learning on the selected speaker pair. This process ensures that when the voice data of the first speaker, one speaker of the speaker pair, is input, the voice data of the second speaker, the other speaker of the speaker pair, is output, and when the voice data of the second speaker is input, the voice data of the first speaker is output (step S310).

[0140] When the voice characteristic conversion model 133 is learned for one speaker pair, the second learning unit 132 determines whether there is an unselected speaker pair that has not been selected among the plurality of speakers for which the voice characteristic conversion model 133 is to be learned (step S320 ).

[0141] In the process of step S320 , when it is determined that there is an unacquired speaker pair (step S320 : Yes), the second learning unit 132 selects one speaker pair from the unselected speaker pairs (step S330 ), and the process proceeds to the process of step S310 .

[0142] In the process of step S320 , when it is determined that there is no unacquired speaker pair (step S320 : No), the speaker identification apparatus 1 ends the voice characteristic conversion model learning process.

[0143] Figure 7 is a flowchart of the speaker identification process.

[0144] Speaker identification processing is processing for identifying the speaker of the speech contained in the speech data. More specifically, speaker identification processing is processing for inputting speech data into the speaker identification model 20 that has been previously trained, and causing the speaker identification model 20 to output speaker identification information.

[0145] The speaker identification process is started, for example, when a user of the speaker identification apparatus 1 performs an operation on the speaker identification apparatus 1 to start the speaker identification process.

[0146] When the speaker identification process starts, the recognition target speech data acquisition unit 40 acquires speech data to be recognized (step S400 ).

[0147] When the voice data is acquired, the third feature value calculation unit 21 calculates a vocalization feature value representing the characteristics of the vocalization included in the voice data based on the acquired voice data (step S410), and inputs the calculated vocalization feature value into the deep neural network 22. The deep neural network 22 then outputs a speaker feature value representing the characteristics of the speaker of the vocalization included in the voice data that served as the source of the calculation of the input vocalization feature value (step S420).

[0148] When the speaker characteristic value is output, the determination unit 23 determines the speaker of the speech included in the speech data acquired by the recognition target speech data acquisition unit 40 based on the output speaker characteristic value (step S430). The determination unit 23 then outputs speaker identification information identifying the determined speaker (step S440).

[0149] When the process of step S440 ends, the speaker identification apparatus 1 ends the speaker identification process.

[0150] <Inspection>

[0151] As described above, the speaker identification device 1 expands the learning data stored in the audio data storage unit 11 for use in learning the speaker identification model 20, without being restricted by utterance content or language. The expanded learning data is then used to learn the speaker identification model 20. Therefore, the speaker identification device 1 can improve the accuracy of speaker identification using the speaker identification model 20. Therefore, the speaker identification device 1 can identify speakers with high accuracy.

[0152] (Supplementary explanation)

[0153] As mentioned above, the speaker recognition device according to the embodiment has been described, but the present disclosure is not limited to this embodiment.

[0154] For example, each processing unit included in the speaker recognition device of the above embodiment is typically implemented as an integrated circuit (LSI). These units may be implemented individually on a single chip, or some or all of them may be implemented on a single chip.

[0155] Furthermore, integrated circuits are not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays) that can be programmed after LSI fabrication, or reconfigurable processors that can reconfigure the connections and settings of circuit cells within the LSI, can also be used.

[0156] Furthermore, the present disclosure can be implemented as a method for learning a speaker recognition model executed by a speaker recognition apparatus according to an embodiment, or can be implemented as a speaker recognition method.

[0157] In addition, in the above-mentioned embodiments, each component is composed of dedicated hardware, but it can also be implemented by executing a software program suitable for each component. Each component can also be implemented by a program execution unit such as a CPU or a processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0158] The division of functional blocks in the block diagram is merely an example. Multiple functional blocks may be implemented as a single functional block, or a single functional block may be divided into multiple blocks, or some functions may be transferred to other functional blocks. Furthermore, a single piece of hardware or software may process the functions of multiple functional blocks with similar functions in parallel or in a time-sharing manner.

[0159] In addition, the order in which each step in the flowchart is executed is exemplified for the purpose of specifically explaining the present disclosure, and may be an order other than the above-described order. In addition, some of the above-described steps may be executed simultaneously (in parallel) with other steps.

[0160] While one or more speaker recognition devices have been described above based on implementations, the present disclosure is not limited to these implementations. Within the scope of the present disclosure, various variations of the present embodiment conceivable by those skilled in the art, as well as combinations of components from various variations, may also be included within the scope of one or more implementations.

[0161] Industrial applicability

[0162] The present disclosure can be widely utilized in devices for recognizing speakers, and the like.

Claims

1. A method for learning a speaker recognition model, wherein the speaker recognition model, when input with voice data, outputs speaker recognition information identifying a speaker uttering the voice contained in the voice data, wherein: generating second speech data of the second speaker by performing speech characteristic conversion processing based on the speech data of the first speaker and the speech data of the second speaker on first speech data of the first speaker, calculating a vocal feature value of the second sound data, determining whether the similarity between the calculated vocal feature quantity of the second voice data and the vocal feature quantity of the second speaker is within a predetermined range, and if the determination is positive, associating speaker identification information identifying the second speaker with the second voice data; and if the determination is negative, associating speaker identification information identifying a third speaker different from the second speaker with the second voice data, The speaker recognition model is trained using the first and second voice data as learning data.

2. The learning method according to claim 1, wherein: The voice characteristic transformation processing includes inputting the first voice data into a voice characteristic transformation model, thereby outputting the second voice data from the voice characteristic transformation model. The voice characteristic transformation model has been pre-learned so that when the voice data of the first speaker is input, the voice data of the second speaker is output.

3. The learning method according to claim 2, wherein: The sound characteristic transformation model includes a deep neural network that takes sound data in WAV format as input and outputs sound data in WAV format.

4. The learning method according to claim 1, wherein: The speaker recognition model includes a deep neural network that takes as input an utterance feature quantity representing a feature of an utterance included in voice data and outputs a speaker feature quantity representing a feature of a speaker.

5. A method for speaker identification, wherein: Voice data is input to the speaker identification model that has been previously trained by the learning method according to claim 1, so that the speaker identification model outputs the speaker identification information.

6. A computer-readable recording medium having a program recorded thereon, the program causing a computer to execute a process for learning a speaker recognition model, the speaker recognition model receiving input of voice data and outputting speaker recognition information identifying a speaker uttering an utterance contained in the voice data, wherein: The processing includes: In a first step, a voice characteristic conversion process based on the voice data of the first speaker and the voice data of the second speaker is performed on the first voice data of the first speaker to generate the second voice data of the second speaker; and In the second step, the first sound data and the second sound data are used as learning data to perform learning processing on the speaker recognition model. The process further comprises the following steps: calculating a vocal feature value of the second sound data; and Determine whether the similarity between the calculated voice feature quantity of the second sound data and the voice feature quantity of the second speaker is within a specified range. If an affirmative determination is made, associate speaker identification information identifying the second speaker with the second sound data. If a negative determination is made, associate speaker identification information identifying a third speaker different from the second speaker with the second sound data.

Citation Information

Patent Citations

  • Frame mapping approach for cross-lingual voice transformation

    US20120253781A1