Speaker recognition method, speaker recognition device, and speaker recognition program
By extracting speaker vectors for partial segments and integrating models to calculate similarity, the method addresses the challenge of low verification accuracy in short utterances, enhancing speaker verification accuracy by considering specific partial segment characteristics.
Patent Information
- Application Number
- JP2022564895
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2020-11-25
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-11-25
AI Technical Summary
Existing speaker verification technologies struggle to accurately quantify speaker characteristics in short utterances due to variations in utterance length, leading to low verification accuracy, especially when speaker characteristics are strongly expressed in specific partial segments of speech.
A method that extracts speaker vectors for each partial segment of a predetermined length from both registered and verification utterances, generating a model to calculate similarity between these segments, integrating a speaker vector extraction model and a speaker similarity calculation sub-model to enhance accuracy.
Enables accurate speaker verification by considering speaker characteristics in partial segments of speech, improving verification accuracy by reflecting specific partial segment characteristics in the speaker vector.
Smart Images

Figure 0007700801000002 
Figure 0007700801000003 
Figure 0007700801000004
Abstract
Description
Technical Field
[0001] The present invention relates to a speaker recognition method, a speaker recognition device, and a speaker recognition program.
Background Art
[0002] In recent years, there has been an expectation for a technology that automatically verifies whether a short utterance is from a person registered in advance. If a speaker can be automatically estimated from a short utterance, for example, in a contact center, it becomes possible to identify a customer from the voice of a call and confirm the identity of the person. Then, since it is no longer necessary to ask for the name, address, customer ID, etc., the call time is reduced, leading to a reduction in operating costs. In interaction with a smart speaker or the like, automatic verification of the speaker using the utterance log becomes possible. Then, it becomes possible to identify a family member from the voice, and it becomes possible to present information and make recommendations according to the speaker.
[0003] For such applications, as an utterance for pre-registering a speaker (hereinafter referred to as a registration utterance), an utterance about several minutes long is used. On the other hand, as an utterance for verifying a speaker (hereinafter referred to as a verification utterance), a short utterance including an arbitrary phrase of about several seconds is used, and a technique called text-independent speaker verification for short utterances is applied.
[0004] In text-independent speaker verification, features such as an x-vector representing speaker characteristics (hereinafter referred to as a speaker vector) indicating that the voice is from the speaker himself / herself expressed in the voice are extracted from the voice, and based on the similarity between speaker vectors, a speaker similarity indicating the identity of the speaker is calculated (see Non-Patent Document 1).
[0005] Conventionally, an x-vector is extracted using a neural network (hereinafter referred to as a speaker vector extraction model). Also, the speaker similarity is quantified using PLDA (Probabilistic Linear Discriminant Analysis), cosine distance, or the like.
[0006] However, when applying the prior art to text-independent speaker verification for short utterances, the difference in utterance length between the registered utterance and the verification utterance is expressed in the speaker vector, making it difficult to accurately quantify the speaker characteristics of the registered utterance and the verification utterance. Therefore, it is known that the verification accuracy decreases.
[0007] Therefore, in the evaluation of speaker similarity, techniques for reducing the variation in speaker similarity due to the difference in utterance length (see Non-Patent Document 2) and techniques for using whether or not the similarity as an audio signal is high for identity determination (see Non-Patent Document 3) have been proposed.
[0008] Note that Non-Patent Document 4 describes the attention mechanism layer in deep learning. Also, Non-Patent Document 5 describes phoneme bottleneck features and the like.
Prior Art Documents
Non-Patent Documents
[0009]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
Non-Patent Document 4
Non-Patent Document 5
Summary of the Invention
Problems to be Solved by the Invention
[0010] However, in the prior art, it has been difficult to perform speaker verification considering the speaker characteristics expressed in the partial segments of speech. That is, even when using the prior art for short speeches, it is impossible to consider the speaker characteristics expressed in specific partial segments of the speech, and the speaker verification accuracy remains low. For example, the characteristics of a cutesy voice may be generated by nasalization in the pronunciation segment of / a / , or the characteristics of a weak voice may be generated by the tongue rising in the pronunciation segment of plosives such as / s / or / t / . Thus, speaker characteristics may be strongly expressed in specific partial segments of the speech. In such cases where the speaker's characteristics are strongly manifested in specific partial segments, in the prior art, since one speaker vector is extracted from the entire speech segment, it is difficult for the characteristics of specific partial segments to be reflected in the speaker vector, and it has been difficult to perform speaker verification considering the speaker characteristics expressed in specific partial segments of the speech.
[0011] The present invention has been made in view of the above, and an object thereof is to perform speaker verification considering the speaker characteristics expressed in partial segments of speech.
Means for Solving the Problem
[0012] In order to solve the above-described problems and achieve the object, a speaker recognition method according to the present invention includes an extraction step of extracting a speaker vector representing the characteristics of a speaker's voice for each partial segment of a predetermined length of an audio signal of speech, and using the speaker vector for each of the partial segments extracted from the audio signal of the speech of a speaker registered in advance and the speaker vector for each of the partial segments extracted from the audio signal of the speech of a speaker to be verified, a learning step of generating by learning a model for calculating the similarity between the audio signal of the speech of the registered speaker and the audio signal of the speech of the speaker to be verified.
Effects of the Invention
[0013] According to the present invention, it becomes possible to perform speaker verification considering the speaker characteristics expressed in partial segments of speech.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0015] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings. Note that the present invention is not limited by this embodiment. Also, in the description of the drawings, the same parts are denoted by the same reference numerals.
[0016] [Outline of Speaker Recognition Device] FIG. 1 is a diagram for explaining the outline of a speaker recognition device. As shown in FIG. 1(a), speaker characteristics are strongly expressed in specific partial intervals rather than the whole utterance. In the example shown in FIG. 1, for example, speaker characteristics are expressed in partial intervals such as "ha" of a nasalized registered utterance, "ka" of a verification utterance, "sou" of a registered utterance with a plosive sound, and "sott" of a verification utterance. In this case, as in the past, it is difficult to say that the speaker vectors extracted from the wholes of registered utterances with different interval lengths and the speaker vectors extracted from the whole of a verification utterance appropriately express speaker characteristics. Therefore, even if such speaker vectors are compared with each other to calculate the similarity, it is difficult to say that it can be used for speaker similarity.
[0017] Therefore, as shown in FIG. 1(b), the speaker recognition device of the present embodiment cuts out a registered utterance and a verification utterance into short partial intervals with a fixed length such as a 1-second width and a 0.5-second shift, and extracts a speaker vector for each partial interval. In this way, it becomes possible to reflect the speaker characteristics expressed for each specific partial interval of the utterance in the speaker vector. The speaker recognition device generates a model (speaker vector extraction model) for extracting a speaker vector by learning.
[0018] Then, as shown in FIG. 1(c), the speaker recognition device compares the speaker vectors of each partial interval of the registered utterance and the speaker vectors of each partial interval of the verification utterance one by one to calculate the respective similarities S. Further, the speaker recognition device generates a model (speaker similarity calculation sub-model) for calculating the speaker similarity y by calculating the weighted sum of the weights α of each similarity S as the speaker similarity y by learning.
[0019] In particular, as shown in FIG. 1(d), the speaker recognition device according to the present embodiment generates, by learning, two models, i.e., the above-described speaker vector extraction model and the speaker similarity calculation sub-model, as an integrated speaker similarity calculation model. Then, the speaker recognition device uses the generated speaker similarity calculation model to output a speaker similarity, for example, 0.5, for an input of a registered utterance and a verification utterance. Further, the speaker recognition device estimates whether the speaker of the registered utterance and the verification utterance match or do not match based on the output speaker similarity. In this way, the speaker recognition device can perform speaker verification considering the speaker characteristics expressed in a partial section of the utterance.
[0020] [First Embodiment] [Configuration of Speaker Recognition Device] FIG. 2 is a schematic diagram illustrating a schematic configuration of the speaker recognition device according to the first embodiment. FIGS. 3 and 4 are diagrams for explaining the processing of the speaker recognition device according to the first embodiment. First, as illustrated in FIG. 2, the speaker recognition device 10 is realized by a general-purpose computer such as a personal computer, and includes an input unit 11, an output unit 12, a communication control unit 13, a storage unit 14, and a control unit 15.
[0021] The input unit 11 is realized by using an input device such as a keyboard or a mouse, and inputs various instruction information such as a processing start to the control unit 15 in response to an input operation by an operator. The output unit 12 is realized by a display device such as a liquid crystal display, a printing device such as a printer, an information communication device, or the like.
[0022] The communication control unit 13 is realized by a NIC (Network Interface Card) or the like, and controls communication between an external device such as a server via a network and the control unit 15. For example, the communication control unit 13 controls communication between the control unit 15 and a management device that manages an audio signal of an utterance.
[0023] The storage unit 14 is implemented by a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. Note that the storage unit 14 may be configured to communicate with the control unit 15 via the communication control unit 13. In the present embodiment, the storage unit 14 stores, for example, a speaker similarity calculation model 14a used for speaker recognition processing described later. Further, the storage unit 14 may store the voice signal of the registered utterance described later.
[0024] The control unit 15 is implemented using a CPU (Central Processing Unit), an NP (Network Processor), an FPGA (Field Programmable Gate Array), etc., and executes a processing program stored in a memory. Thereby, as illustrated in FIG. 2, the control unit 15 functions as an acoustic feature extraction unit 15a, a speaker vector extraction unit 15b, a learning unit 15c, a calculation unit 15d, and an estimation unit 15e. Note that these functional units may be implemented by different hardware. For example, the learning unit 15c may be implemented as a learning device, and the calculation unit 15d and the estimation unit 15e may be implemented as an estimation device. Further, the control unit 15 may include other functional units.
[0025] The acoustic feature extraction unit 15a extracts the acoustic features of the voice signal of the utterance. For example, the acoustic feature extraction unit 15a receives the input of the voice signal of the registered utterance and the voice signal of the verification utterance via the input unit 11 or via the communication control unit 13 from a management device that manages the voice signal of the utterance. Further, the acoustic feature extraction unit 15a extracts acoustic features for each partial section (short-time window) of the voice signal of the utterance, and outputs an acoustic feature sequence in which vectors of acoustic features (speaker vectors) are arranged in time series order. The acoustic features include, for example, information including any one or more of a power spectrum, a logarithmic mel filter bank, an MFCC (Mel Frequency Cepstral Coefficient), a fundamental frequency, a logarithmic power, and a first derivative or a second derivative thereof. Alternatively, the acoustic feature extraction unit 15a may use the voice signal as it is without extracting the acoustic feature sequence.
[0026] The speaker vector extraction unit 15b extracts a speaker vector representing the characteristics of the speaker's voice for each predetermined-length partial section of the voice signal of the utterance. Specifically, the speaker vector extraction unit 15b first obtains, from the acoustic feature extraction unit 15a, the voice signal or acoustic feature sequence of a registered utterance that is the utterance of a pre-registered speaker, and the voice signal or acoustic feature sequence of a collation utterance that is the utterance of the speaker to be collated. In the following description, the "voice signal or acoustic feature sequence" may be simply referred to as a voice signal.
[0027] Also, as shown in FIG. 4, the speaker vector extraction unit 15b cuts out each of the acquired voice signals of the registered speaker and the voice signal of the collation speaker into short partial sections of a fixed length such as a 1-second width and a 0.5-second shift, and extracts a speaker vector from each partial section. As shown in FIG. 4, the speaker vector extraction unit 15b extracts a speaker vector from each partial section of the voice signal of the utterance using the speaker vector extraction model 14b.
[0028] Note that the speaker vector extraction unit 15b may be included in the learning unit 15c and the calculation unit 15d described later. For example, FIGS. 3 and 8 described later show examples in which the learning unit 15c and the calculation unit 15d perform the processing of the speaker vector extraction unit 15b. By including the processing of the speaker vector extraction unit 15b in the learning unit 15c, as will be described later, it becomes possible to integrally learn the speaker vector extraction model 14b and the speaker similarity calculation sub-model 14c.
[0029] Return to the description of FIG. 2. The learning unit 15c generates, by learning, a speaker similarity calculation sub-model 14c that calculates the similarity between the voice signal of the utterance of a pre-registered speaker and the voice signal of the utterance of a speaker to be verified, using the speaker vectors for each partial section extracted from the voice signal of the utterance of the pre-registered speaker and the speaker vectors for each partial section extracted from the voice signal of the utterance of the speaker to be verified. That is, as shown in FIG. 3, the learning unit 15c performs learning of a speaker similarity calculation model 14a including the speaker similarity calculation sub-model 14c, using the speaker vectors of the registered utterance and the verification utterance extracted by the speaker vector extraction unit 15b and the speaker match / mismatch information indicating whether the speaker of the registered utterance and the speaker of the verification utterance match or not.
[0030] Specifically, as shown in FIG. 4, the learning unit 15c generates a speaker similarity calculation sub-model 14c represented by the weighted sum of the similarities of each partial section of the speaker vector of the utterance of the registered speaker and the speaker vector of each partial section of the utterance of the speaker to be verified.
[0031] That is, the learning unit 15c calculates the similarity S for each by comparing the speaker vectors of each partial section of the voice signal of the registered utterance and the speaker vectors of each partial section of the voice signal of the verification speaker on a one-to-one basis. Further, the learning unit 15c generates, by learning, a speaker similarity calculation sub-model 14c that calculates the speaker similarity y, which is the weighted sum of the weights α of each similarity S, using, for example, the speaker match / mismatch information represented by 1 / 0. Here, the speaker similarity y is expressed as the following formula (1).
[0032]
Equation
[0033] For example, the attention mechanism layer shown in FIG. 4 combines the speaker vectors of each partial section of the voice signal of the registered utterance and the voice signal of the verification utterance in a one-to-one manner. For each pair, the similarity S between the speaker vectors and the weight α of each similarity are calculated, and a weighted sum is performed. Further, the pooling layer averages the feature vectors representing the similarity of the registered utterance for each partial section of the verification utterance output from the attention mechanism layer, and the fully connected layer and the activation function convert it into a scalar value, thereby calculating the speaker similarity y.
[0034] Further, the learning unit 15c generates, by learning, a speaker vector extraction model 14b from which the speaker vector extraction unit 15b extracts a speaker vector. That is, as shown in FIGS. 3 and 4, the learning unit 15c of the present embodiment generates, by learning, the speaker similarity calculation sub-model 14c and the speaker vector extraction model 14b as an integrated speaker similarity calculation model 14a.
[0035] Specifically, the learning unit 15c optimizes the speaker similarity calculation model 14a using the speaker similarity output from the speaker similarity calculation model 14a and the speaker match / mismatch information. That is, the learning unit 15c cuts out the voice signal of the registered utterance and the partial section of the voice signal of the verification utterance, and for the speaker vector for each partial section extracted using the speaker vector extraction model 14b and the speaker similarity calculated using the speaker similarity calculation sub-model 14c, the speaker vector extraction model 14b and the speaker similarity calculation sub-model 14c are optimized. The learning unit 15c optimizes the speaker vector extraction model 14b and the speaker similarity calculation sub-model 14c so that the speaker similarity output when the speaker of the input registered utterance matches the speaker of the verification utterance is large, and the speaker similarity output when they do not match is small. For example, the learning unit 15c defines cross-entropy error or the like as a loss function, and updates the model parameters of the speaker vector extraction model 14b and the speaker similarity calculation sub-model 14c so that the loss function becomes small using the stochastic gradient descent method.
[0036] As a result, a speaker vector extraction model 14b is generated that can more appropriately extract speaker characteristics for each sub-interval. For example, a speaker vector extraction model 14b is generated that reflects characteristics such that the pronunciation patterns of / s / and / t / are easily quantifiable as speaker vectors, while a mora is difficult to be quantified into a speaker vector. Also, a speaker similarity calculation sub-model 14c is generated that can accurately estimate the similarity S and its weight α for each pair of the sub-intervals of the registered utterance and the comparison utterance. For example, a speaker similarity calculation sub-model 14c is generated in which the weight of the similarity between "sou" of the registered utterance and "sott" of the comparison utterance illustrated in FIG. 1 is high, and the weights of the similarities between other sub-intervals are low.
[0037] Returning to the description of FIG. 2, the calculation unit 15d calculates the similarity between the voice signal of the utterance of a pre-registered speaker and the voice signal of the utterance of the speaker to be compared, using the generated speaker similarity calculation model 14a. Specifically, the calculation unit 15d inputs the speaker vectors of the sub-intervals of the voice signal of the registered utterance and the speaker vectors of the sub-intervals of the voice signal of the comparison speaker, which are extracted by the speaker vector extraction unit 15b using the speaker vector extraction model 14b, into the speaker similarity calculation sub-model 14c, and outputs the speaker similarity. Note that, as shown in FIG. 3, the voice signal of the registered utterance used by the calculation unit 15d does not have to be the same as the voice signal of the registered utterance used by the learning unit 15c, and may be a different voice signal.
[0038] The estimation unit 15e estimates whether or not the speakers of the utterance of the pre-registered speaker and the utterance of the speaker to be compared match, using the calculated similarity. Specifically, as shown in FIG. 3, for example, when the calculated speaker similarity is equal to or greater than a predetermined threshold, the estimation unit 15e estimates that the speakers of the registered utterance and the comparison speaker match, and outputs speaker match / mismatch information indicating the match. Also, when the speaker similarity is less than the predetermined threshold, the estimation unit 15e estimates that the speakers of the registered utterance and the comparison speaker do not match, and outputs speaker match / mismatch information indicating the mismatch.
[0039] [Speaker recognition processing] Next, the speaker recognition process by the speaker recognition device 10 will be described. FIGS. 5 and 6 are flowcharts showing the speaker recognition process procedure. The speaker recognition process of the present embodiment includes a learning process and an estimation process. First, FIG. 5 shows the learning process procedure. The flowchart in FIG. 5 starts at the timing when an input instructing the start of the learning process is given, for example.
[0040] First, the speaker vector extraction unit 15b acquires the voice signal of the registered utterance and the voice signal of the verification utterance from the acoustic feature extraction unit 15a, cuts out each voice signal for each short partial section of a predetermined length, and uses the speaker vector extraction model 14b to extract a speaker vector from each partial section (step S1).
[0041] Next, the learning unit 15c uses the speaker vectors for each partial section extracted from the voice signal of the registered utterance and the speaker vectors for each partial section extracted from the voice signal of the verification utterance to generate, by learning, a speaker similarity calculation sub-model 14c that calculates the similarity between the voice signal of the registered utterance and the voice signal of the verification utterance (step S2).
[0042] Specifically, the learning unit 15c generates, by learning, a speaker vector extraction model 14b for the speaker vector extraction unit 15b to extract a speaker vector. Also, the learning unit 15c compares the speaker vectors of each partial section of the voice signal of the registered utterance with the speaker vectors of each partial section of the voice signal of the verification speaker one by one to calculate the similarity S for each. Further, the learning unit 15c uses the speaker match / mismatch information to generate, by learning, a speaker similarity calculation sub-model 14c that calculates the speaker similarity y, which is the weighted sum of the weights α of each similarity S.
[0043] That is, the learning unit 15c uses the speaker similarity calculation sub-model 14c and the speaker vector extraction model 14b as an integrated speaker similarity calculation model 14a, and uses the speaker similarity output from the speaker similarity calculation model 14a and the speaker match / mismatch information to optimize the speaker similarity calculation model 14a. Thereby, a series of learning processes is completed.
[0044] Next, FIG. 6 shows the estimation processing procedure. The flowchart in FIG. 6 starts, for example, at the timing when an input instructing the start of the estimation processing is received.
[0045] First, the speaker vector extraction unit 15b acquires the voice signal of the registered utterance and the voice signal of the verification utterance from the acoustic feature extraction unit 15a, cuts out each voice signal for short subintervals of a predetermined length, and extracts speaker vectors from each subinterval using the speaker vector extraction model 14b generated by learning (step S1).
[0046] Next, the calculation unit 15d calculates the similarity between the voice signal of the registered utterance and the voice signal of the verification utterance using the generated speaker similarity calculation model 14a (step S3). Specifically, the calculation unit 15d inputs the speaker vector of the subinterval of the voice signal of the registered utterance and the speaker vector of the subinterval of the voice signal of the verification speaker into the speaker similarity calculation submodel 14c and outputs the speaker similarity.
[0047] Also, the estimation unit 15e estimates whether the speakers of the registered utterance and the verification utterance to be verified match or not using the calculated speaker similarity (step S4), and outputs speaker match / mismatch information. Thereby, a series of estimation processing is completed.
[0048] [Second Embodiment] The speaker recognition device 10 is not limited to the above embodiment. For example, the learning unit 15c may further generate the speaker similarity calculation model 14a by learning using the phoneme sequence of the utterance. Hereinafter, the speaker recognition device 10 of this second embodiment will be described with reference to FIGS. 7 to 9. Note that only the differences from the speaker recognition processing of the speaker recognition device 10 of the above first embodiment will be described, and the description of the common points will be omitted.
[0049] FIG. 7 is a schematic diagram illustrating the schematic configuration of the speaker recognition apparatus according to the second embodiment. FIGS. 8 and 9 are diagrams for explaining the processing of the speaker recognition apparatus according to the second embodiment. First, as shown in FIG. 7, the speaker recognition apparatus 10 of the present embodiment is different from the speaker recognition apparatus 10 of the first embodiment described above in that it has a phoneme discrimination model 14d and a recognition unit 15f.
[0050] Specifically, the speaker recognition apparatus 10 of the present embodiment calculates the speaker similarity using further the phonetic information of the registered utterance and the collated utterance. Here, the phonetic information is, for example, the phoneme sequence of the utterance. Alternatively, the phonetic information may be a phoneme posterior probability sequence or a phoneme bottleneck feature output as a latent variable.
[0051] In the speaker recognition apparatus 10 of the present embodiment, as shown in FIG. 8, the recognition unit 15f outputs the phoneme sequence of the input utterance using a pre-learned phoneme discrimination model 14d. Further, as shown in FIG. 9, the speaker vector extraction unit 15b cuts out the utterance phoneme sequence into short partial intervals of a predetermined length such as a 1-second width and a 0.5-second shift, and extracts the speaker vector for each partial interval using the speaker vector extraction model 14b.
[0052] In this case, in addition to the speaker vectors of each partial interval of the registered utterance voice signal and the speaker vectors of each partial interval of the collated speaker's voice signal, the learning unit 15c further uses the speaker vectors of each partial interval of the phoneme sequence of the registered utterance and the speaker vectors of each partial interval of the voice sequence of the collated utterance. Thereby, the learning unit 15c generates a speaker similarity calculation model 14a' considering the phonetic information by learning.
[0053] Further, similar to the first embodiment described above, the learning unit 15c of the present embodiment generates, by learning, the speaker similarity calculation sub-model 14c and the speaker vector extraction model 14b as an integrated speaker similarity calculation model 14a'.
[0054] Specifically, as shown in FIG. 8, the learning unit 15c receives the speaker vectors for each partial section extracted using the speaker vector extraction model 14b from the voice signal of the registered utterance and the phoneme sequence of the registered utterance, the voice signal of the verification utterance, and the voice sequence of the verification utterance, as well as the speaker match / mismatch information. Then, as shown in FIG. 9, the learning unit 15c optimizes the speaker vector extraction model 14b and the speaker similarity calculation sub-model 14c using the speaker similarity calculated using the speaker similarity calculation sub-model 14c and the speaker match / mismatch information.
[0055] As a result, the speaker recognition device 10 can construct a speaker similarity calculation model 14a' that takes into account phonetic information. Therefore, the speaker recognition device 10 can calculate the speaker similarity with higher accuracy, and can accurately estimate whether the speakers match when verifying the registered utterance and the verification utterance.
[0056] As described above, in the speaker recognition device 10 of the present embodiment, the speaker vector extraction unit 15b extracts a speaker vector representing the characteristics of the speaker's voice for each partial section of a predetermined length of the voice signal of the utterance. Further, the learning unit 15c uses the speaker vectors for each partial section extracted from the registered utterance, which is the voice signal of the utterance of a pre-registered speaker, and the speaker vectors for each partial section extracted from the verification utterance, which is the voice signal of the utterance of the speaker to be verified, to generate, by learning, a speaker similarity calculation sub-model 14c that calculates the similarity between the voice signal of the registered utterance and the voice signal of the verification utterance.
[0057] As a result, it becomes possible to perform speaker verification considering the speaker characteristics expressed in the partial sections of the utterance. Therefore, it becomes possible to accurately estimate whether the speakers of the utterance of the speaker registered with high accuracy and the utterance of the speaker to be verified match.
[0058] Further, the learning unit 15c generates a speaker similarity calculation sub-model 14c represented by the weighted sum of the similarities between the speaker vectors of each partial section of the registered utterance and the speaker vectors of each partial section of the verification utterance. As a result, it becomes possible to calculate the speaker similarity with high accuracy.
[0059] Further, the learning unit 15c generates, by learning, a speaker vector extraction model 14b for the speaker vector extraction unit 15b to extract a speaker vector. That is, the learning unit 15c generates, by learning, the speaker similarity calculation sub-model 14c and the speaker vector extraction model 14b as an integrated speaker similarity calculation model 14a. Thereby, a speaker vector extraction model 14b that can more appropriately extract the speaker characteristics for each partial section, and a speaker similarity calculation sub-model 14c that can accurately estimate the similarity S and its weight α for each pair of the partial section of the registered utterance and the partial section of the collation utterance are efficiently generated.
[0060] Also, in the speaker recognition device 10, the calculation unit 15d calculates the speaker similarity between the voice signal of the utterance of a pre-registered speaker and the voice signal of the collation utterance to be collated, using the generated speaker similarity calculation model 14a. Further, the estimation unit 15e estimates whether or not the speakers of the utterance of the registered speaker and the utterance of the speaker to be collated match, using the calculated speaker similarity. Thereby, it becomes possible to accurately estimate whether or not the speakers of the utterance of the registered speaker and the utterance of the speaker to be collated match.
[0061] Further, the learning unit 15c further generates, by learning, a speaker similarity calculation sub-model 14c' using the phoneme sequence of the utterance. Thereby, the speaker recognition device 10 can calculate the speaker similarity with higher accuracy, and when collating the registered utterance and the collation utterance, it becomes possible to accurately estimate whether or not the speakers match.
[0062] [Program] It is also possible to create a program that describes the processing executed by the speaker recognition device 10 according to the above embodiment in a language executable by a computer. As one embodiment, the speaker recognition device 10 can be implemented by installing a speaker recognition program that executes the above speaker recognition processing as package software or online software on a desired computer. For example, by causing the information processing device to execute the above speaker recognition program, the information processing device can be made to function as the speaker recognition device 10. In addition, other information processing devices include mobile communication terminals such as smartphones, mobile phones, and PHS (Personal Handyphone System), and slate terminals such as PDAs (Personal Digital Assistants). Further, the functions of the speaker recognition device 10 may be implemented on a cloud server.
[0063] FIG. 12 is a diagram showing an example of a computer that executes a speaker recognition program. The computer 1000 includes, for example, a memory 1010, a CPU 1020, a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0064] The memory 1010 includes a ROM (Read Only Memory) 1011 and a RAM 1012. The ROM 1011 stores a boot program such as a BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1031. The disk drive interface 1040 is connected to the disk drive 1041. A removable storage medium such as a magnetic disk or an optical disk is inserted into the disk drive 1041. For example, a mouse 1051 and a keyboard 1052 are connected to the serial port interface 1050. For example, a display 1061 is connected to the video adapter 1060.
[0065] Here, the hard disk drive 1031 stores, for example, an OS 1091, an application program 1092, a program module 1093, and program data 1094. Each piece of information described in the above embodiment is stored, for example, in the hard disk drive 1031 or the memory 1010.
[0066] Also, the speaker recognition program is stored in the hard disk drive 1031 as a program module 1093 in which, for example, instructions executed by the computer 1000 are described. Specifically, a program module 1093 in which each process executed by the speaker recognition device 10 described in the above embodiment is described is stored in the hard disk drive 1031.
[0067] Also, data used for information processing by the speaker recognition program is stored, for example, in the hard disk drive 1031 as program data 1094. Then, the CPU 1020 reads out the program module 1093 and the program data 1094 stored in the hard disk drive 1031 to the RAM 1012 as necessary, and executes each of the above-described procedures.
[0068] Note that the program module 1093 and the program data 1094 related to the speaker recognition program are not limited to being stored in the hard disk drive 1031. For example, they may be stored in a removable storage medium and read out by the CPU 1020 via a disk drive 1041 or the like. Alternatively, the program module 1093 and the program data 1094 related to the speaker recognition program may be stored in another computer connected via a network such as a LAN (Local Area Network) or a WAN (Wide Area Network), and read out by the CPU 1020 via the network interface 1070.
[0069] The embodiments to which the invention made by the present inventor has been applied have been described above. However, the present invention is not limited by the description and drawings that form a part of the disclosure of the present invention according to the present embodiment. That is, all other embodiments, examples, operation techniques, etc. made by those skilled in the art based on the present embodiment are included in the scope of the present invention.
Explanation of Signs
[0070] 10 Speaker recognition device 11 Input unit 12 Output unit 13 Communication control unit 14 Storage unit 14a Speaker similarity calculation model 14b Speaker vector extraction model 14c Speaker similarity calculation sub-model 14d Phoneme discrimination model 15 Control unit 15a Acoustic feature extraction unit 15b Speaker vector extraction unit 15c Learning unit 15d Calculation unit 15e Estimation unit 15f Recognition unit
Claims
1. A speaker recognition method executed by a speaker recognition device, comprising: an extraction step of extracting a speaker vector representing the characteristics of a speaker's voice for each predetermined-length partial section of an utterance voice signal; a learning step of generating, by learning, a model for calculating the similarity between the voice signal of the utterance of the registered speaker and the voice signal of the utterance of the speaker to be verified, using the speaker vectors for each of the partial sections extracted from the voice signal of the utterance of the registered speaker and the speaker vectors for each of the partial sections extracted from the voice signal of the utterance of the speaker to be verified; wherein the learning step generates the model that calculates a speaker similarity represented by a weighted sum of the similarities between the speaker vectors of each partial section of the utterance of the registered speaker and the speaker vectors of each partial section of the utterance of the speaker to be verified; A speaker recognition method characterized by the above.
2. The speaker recognition method according to claim 1, wherein the learning step further generates, by learning, a model by which the extraction step extracts the speaker vector.
3. a calculation step of calculating the similarity between the voice signal of the utterance of a registered speaker and the voice signal of the utterance of a speaker to be verified, using the generated model; an estimation step of estimating whether or not the speakers of the utterance of the registered speaker and the utterance of the speaker to be verified match, using the calculated similarity; The speaker recognition method according to claim 1, further comprising the above.
4. The speaker recognition method according to claim 1, wherein the learning step further generates, by learning, the model using a phoneme sequence of an utterance.
5. an extraction unit that extracts a speaker vector representing the characteristics of a speaker's voice for each predetermined-length partial section of an utterance voice signal; a learning unit that generates, by learning, a model for calculating the similarity between the voice signal of the utterance of a registered speaker and the voice signal of the utterance of a speaker to be verified, using the speaker vectors for each of the partial sections extracted from the voice signal of the utterance of the registered speaker and the speaker vectors for each of the partial sections extracted from the voice signal of the utterance of the speaker to be verified; wherein the learning unit generates the model that calculates a speaker similarity represented by a weighted sum of the similarities between the speaker vectors of each partial section of the utterance of the registered speaker and the speaker vectors of each partial section of the utterance of the speaker to be verified; A speaker recognition device characterized by the above.
6. An extraction step of extracting a speaker vector representing the characteristics of the speaker's voice for each predetermined-length partial section of the speech audio signal, A speaker recognition program for causing a computer to execute a learning step of generating, by learning, a model for calculating the similarity between the speech audio signal of the registered speaker and the speech audio signal of the speaker to be verified, using the speaker vectors for each of the partial sections extracted from the speech audio signal of the registered speaker's speech and the speaker vectors for each of the partial sections extracted from the speech audio signal of the speaker to be verified, wherein the learning step generates the model that calculates a speaker similarity represented by a weighted sum of the similarities of each partial section of the speech of the registered speaker and each partial section of the speech of the speaker to be verified, Speaker recognition program.
Citation Information
Patent Citations
Method and device for verifying speaker authentication, and speaker authentication system
JP2009151305A
Speaker recognition device, speaker recognition method, and speaker recognition program
JP2014048534A
Voice processing apparatus, voice processing method, and voice processing program
JP2017058483A
Speaker-likeness evaluation device, speaker identification device, speaker collation device, speaker-likeness evaluation method, and program
JP2017097188A