Voice authentication device and voice authentication method
The voice authentication system addresses accuracy issues by determining sound pickup conditions and selecting appropriate models, enhancing speaker recognition despite varying noise and conditions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
- Filing Date
- 2023-03-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing voiceprint authentication systems face accuracy issues due to changes in ambient noise and sound collection conditions, leading to inconsistent feature extraction and authentication results.
A voice authentication system that acquires audio data, detects speech segments, extracts features, determines sound pickup conditions, and selects a similarity calculation model based on these conditions to authenticate speakers accurately.
The system effectively suppresses the decline in speaker recognition accuracy caused by ambient noise changes by selecting appropriate models for varying sound conditions, ensuring accurate authentication.
Smart Images

Figure 0007867188000002 
Figure 0007867188000003 
Figure 0007867188000004
Abstract
Description
[Technical Field]
[0001] This disclosure relates to a voice authentication device and a voice authentication method. [Background technology]
[0002] Patent Document 1 discloses a speech recognition device for recognizing a subject's voice. The speech recognition device stores multiple action noise models, each corresponding to a plurality of actions, and associates them with each of the plurality of actions. It detects input audio containing the subject's voice, identifies the subject's actions, and reads out the action noise model corresponding to the action identified by the action identification means. The speech recognition device reads out an environmental noise model corresponding to the subject's current location, synthesizes the environmental noise model with the read-out action noise model, and uses the synthesized noise superposition model to recognize the subject's voice contained in the detected input audio. [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2008-250059 [Overview of the project] [Problems that the invention aims to solve]
[0004] However, in Patent Document 1, it was necessary to collect operational noise generated at each of the multiple operations and ambient noise at each of the multiple locations where speech recognition could be performed in advance, which was very time-consuming. Furthermore, in voiceprint authentication, the feature quantities that indicate the individuality of the authentication target (person) extracted from the audio signal change depending on the noise contained in the audio signal and the sound collection conditions of the sound collection equipment used to collect the audio signal. Therefore, when performing voiceprint authentication using the above-mentioned speech recognition device, if the sound collection conditions of the pre-registered audio signal and the sound collection conditions of the audio signal collected at the time of voiceprint authentication are different, the feature quantities extracted from each audio signal may not indicate the individuality of the same person, potentially leading to a decrease in voiceprint authentication accuracy.
[0005] This disclosure was devised in view of the conventional circumstances described above and aims to provide a voice authentication device and voice authentication method that suppress the decline in speaker authentication accuracy caused by changes in ambient noise.
[0006] This disclosure includes an acquisition unit that acquires audio data, a detection unit that detects a speech segment spoken by a speaker from the audio data, and an extraction unit that extracts the speaker's speech features from the detected speech segment. A determination unit that determines the sound pickup conditions for the speech corresponding to the speech features based on the speech features, The extracted speech features of the speaker and the pre-registered multiple The system comprises: a selection unit that selects a first similarity calculation model from among multiple similarity calculation models to be used for speaker authentication based on the speech features of the registered speaker; and an authentication unit that authenticates the speaker by comparing the speech features of the registered speaker with the speech features of the registered speaker using the selected first similarity calculation model. The selection unit selects the first similarity calculation model based on the sound acquisition conditions corresponding to the speech features of the speaker and the sound acquisition conditions corresponding to the speech features of the registered speaker. To provide a voice authentication device.
[0007] Furthermore, this disclosure relates to a voice authentication method performed by a terminal device, comprising: acquiring voice data; detecting a speech segment spoken by a speaker from the voice data; and extracting the speaker's speech features from the detected speech segment. Based on the aforementioned speech features, the conditions for capturing the speech corresponding to the aforementioned speech features are determined. Extracted speech features of the aforementioned speaker Corresponding sound pickup conditions And, pre-registered multiple Speech features of registered speakers Corresponding sound pickup conditions Based on this, a first similarity calculation model is selected from among a plurality of similarity calculation models to be used for speaker authentication, and the selected first similarity calculation model is used to compare the utterance features of the speaker with the utterance features of the registered speaker to authenticate the speaker. [Effects of the Invention]
[0008] According to this disclosure, it is possible to suppress the decline in speaker recognition accuracy caused by changes in ambient noise. [Brief explanation of the drawing]
[0009] [Figure 1] Block diagram showing an example of the internal configuration of the voice authentication system according to the embodiment. [Figure 2] Figure for explaining each process performed by the processor of the terminal device in the embodiment [Figure 3] Flowchart showing an example of the operation procedure of the terminal device in the embodiment [Figure 4] Flowchart showing an example of the procedure for determining the sound collection condition of the terminal device in the embodiment [Figure 5] Figure for explaining an example of the determination of the sound collection condition and an example of the calculation of the reliability [Figure 6] Flowchart showing an example of the speaker authentication procedure of the terminal device in the embodiment [Figure 7] Figure for explaining an example of the correspondence list when the sound collection condition estimation result at the time of voice registration is the same as the sound collection condition estimation result at the time of voice authentication [Figure 8] Figure for explaining an example of the correspondence list when the noise type at the time of voice registration is different from the noise type at the time of voice authentication [Figure 9] Figure for explaining a specific example of the correspondence list [[ID=**********]]
Embodiment for Carrying Out the Invention
[0010] Hereinafter, embodiments specifically disclosing the voice authentication device and the voice authentication method according to the present disclosure will be described in detail with reference to the drawings as appropriate. However, a more detailed description than necessary may be omitted. For example, detailed descriptions of already well-known matters and duplicate descriptions for substantially the same configurations may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art. Note that the attached drawings and the following description are provided for those skilled in the art to fully understand the present disclosure, and it is not intended to limit the subject matter described in the claims by these.
[0011] First, referring to FIGS. 1 and 2, the voice authentication system 100 according to the embodiment will be described. FIG. 1 is a block diagram showing an example of the internal configuration of the voice authentication system 100 according to the embodiment. FIG. 2 is a figure for explaining each process performed by the processor 11 of the terminal device P1 in the embodiment.
[0012] The voice authentication system 100 includes a terminal device P1 as an example of a voice authentication device and a monitor MN. Note that the voice authentication system 100 may be configured to include a microphone MK or a monitor MN.
[0013] The microphone MK picks up the speech voice of the speaker US for voice registration in the terminal device P1 in advance. The microphone MK converts the picked-up speech voice of the speaker US into a voice signal or voice data to be registered in the terminal device P1. The microphone MK transmits the converted voice signal or voice data to the processor 11 via the communication unit 10.
[0014] Also, the microphone MK picks up the speech voice of the speaker US used for speaker authentication. The microphone MK converts the picked-up speech voice of the speaker US into a voice signal or voice data. The microphone MK transmits the converted voice signal or voice data to the processor 11 via the communication unit 10.
[0015] In the following description, for the sake of clarity, the voice data for voice registration or the voice data already registered in the terminal device P1 is denoted as "registered voice data", and the voice data for voice authentication is denoted as "authentication voice data" and distinguished.
[0016] Note that the microphone MK may be a microphone included in a predetermined device such as a Personal Computer (hereinafter referred to as "PC"), a notebook PC, a smartphone, a tablet terminal, etc. Also, the microphone MK may transmit voice signals or voice data to the terminal device P1 by wireless communication via a network (not shown).
[0017] Terminal device P1 is implemented by, for example, a PC, notebook PC, smartphone, tablet terminal, etc., and performs voice registration processing using registered voice data of speaker US and speaker authentication processing using authenticated voice data. It includes a communication unit 10, a processor 11, memory 12, a feature extraction model database DB1, a registered speaker database DB2, a similarity calculation model database DB3, and a learning database DB4 based on sound pickup conditions.
[0018] As an example of an acquisition unit, the communication unit 10 is connected to the microphone MK and the monitor MN via wired or wireless communication, enabling data transmission and reception between them. Wireless communication here refers to, for example, short-range wireless communication such as Bluetooth®, NFC®, or communication via a wireless Local Area Network (LAN) such as Wi-Fi®.
[0019] The communication unit 10 may also transmit and receive data with the microphone MK via an interface such as Universal Serial Bus (USB). Furthermore, the communication unit 10 may also transmit and receive data with the monitor MN via an interface such as High-Definition Multimedia Interface (HDMI, registered trademark).
[0020] The processor 11 is configured using, for example, a Central Processing Unit (CPU) or a Field Programmable Gate Array (FPGA), and works in cooperation with the memory 12 to perform various processes and controls. Specifically, the processor 11 refers to the programs and data held in the memory 12 and executes those programs to realize the functions of each unit, such as the feature extraction unit 111, the sound pickup condition determination unit 112, the speaker registration unit 113, the similarity calculation model selection unit 114, the confidence calculation unit 115, and the authentication unit 116.
[0021] When registering speaker US's voice, the processor 11 implements the functions of the feature extraction unit 111, the speaker registration unit 113, and the sound pickup condition determination unit 112, thereby executing the process of newly registering (storing) speaker US in the registered speaker database DB2.
[0022] Furthermore, when authenticating the voice of speaker US, the processor 11 executes speaker authentication by realizing the functions of the feature extraction unit 111, the sound pickup condition determination unit 112, the similarity calculation model selection unit 114, the confidence calculation unit 115, and the authentication unit 116.
[0023] The feature extraction unit 111, as an example of a detection and extraction unit, acquires speaker US voice data (registered voice data or authenticated voice data) transmitted from microphone MK. Based on the control commands associated with the voice data, the feature extraction unit 111 performs voice registration processing or voice authentication processing.
[0024] The feature extraction unit 111 detects the utterance segments spoken by speaker US from the registered voice data during voice registration. The feature extraction unit 111 extracts features that indicate the individuality of speaker US from the detected utterance segments and outputs them to the sound acquisition condition determination unit 112 and the speaker registration unit 113.
[0025] Furthermore, during voice authentication, the feature extraction unit 111 extracts speaker US features from the authenticated voice data and outputs them to the sound pickup condition determination unit 112 and the authentication unit 116, respectively.
[0026] As an example of a determination unit, the sound pickup condition determination unit 112 determines the sound pickup conditions under which the audio data from which these features were extracted (i.e., the speech of speaker US) was picked up, based on the speaker US feature quantities output from the feature quantity extraction unit 111. When registering speech, the sound pickup condition determination unit 112 outputs information on the sound pickup conditions of speaker US feature quantities to the speaker registration unit 113. Also, when performing speech authentication, the sound pickup condition determination unit 112 outputs information on the sound pickup conditions of speaker US feature quantities to the similarity calculation model selection unit 114.
[0027] The sound acquisition conditions referred to here include the sound acquisition device used to capture the speech of the US speaker or registered speaker, the language, gender, age, and noise type of the US speaker or registered speaker, and the noise type included in the feature vectors. Examples of sound acquisition devices include microphones, telephones, and headsets.
[0028] Noise is noise picked up due to the environment (background) at the time of sound recording, such as surrounding conversations, music, vehicle noise, wind noise, etc. The noise type indicates the environment (location) and position where the noise occurs, such as store noise, outdoor wind noise, store music, or train station premises. The noise type may also include information about the time of day, such as early morning, daytime, or nighttime.
[0029] The speaker registration unit 113 acquires the speaker US feature quantities output from the feature extraction unit 111 and the speaker information of the speaker US associated with the registered speech data. The speaker registration unit 113 also acquires the sound pickup condition information output from the sound pickup condition determination unit 112. The speaker registration unit 113 associates the speaker US feature quantities, the speaker information of the speaker US, and the sound pickup condition information and registers them in the registered speaker database DB2.
[0030] Speaker information may be extracted from registered voice data using speech recognition, or it may be obtained from a device owned by the speaker (e.g., a PC, laptop, smartphone, or tablet). The speaker information referred to here includes, for example, identification information that can identify the speaker, the speaker's name, and speaker identification (ID).
[0031] As an example of a selection unit, the similarity calculation model selection unit 114 acquires information on the sound acquisition conditions of the speaker US feature quantities output from the sound acquisition condition determination unit 112, and information on the sound acquisition conditions of the feature quantities of each of the multiple registered speakers registered in the registered speaker database DB2. Based on the acquired information on the sound acquisition conditions of the speaker US feature quantities and the information on the sound acquisition conditions of each of the multiple registered speakers, the similarity calculation model selection unit 114 selects a similarity calculation model (an example of a similarity calculation model) to be used in the process of calculating the similarity between the speaker US feature quantities and the feature quantities of any one of the registered speakers.
[0032] The similarity calculation model selection unit 114 refers to correspondence lists LST, LST1, and LST2 (see Figures 7, 8, and 9), which associate the acquired speaker US feature information with the respective feature information of multiple registered speakers and the selected similarity calculation model, to select one of the multiple selection models (first similarity calculation models) and output it to the confidence calculation unit 115 and the authentication unit 116, respectively.
[0033] As an example of a confidence calculation unit, the confidence calculation unit 115 calculates (evaluates) a confidence score that indicates the likelihood of the speaker US identification result based on the similarity calculated by the authentication unit 116. The confidence calculation unit 115 calculates the confidence score based on the distance between the model training data distribution of the similarity calculation model used in the similarity calculation process by the authentication unit 116 and the speaker US features. The confidence calculation unit 115 outputs the calculated confidence score information to the authentication unit 116.
[0034] As an example of a calculation unit, the authentication unit 116 acquires the speaker US feature output from the feature extraction unit 111 and acquires the feature of each of the multiple registered speakers registered in the registered speaker database DB2. The authentication unit 116 also acquires the selected model from the correspondence lists LST, LST1, and LST2 output from the similarity calculation model selection unit 114.
[0035] The authentication unit 116 uses a similarity calculation model based on the correspondence lists LST, LST1, and LST2 to calculate the similarity between the features of each of the multiple registered speakers and the features of speaker US. Based on the calculated similarity, the authentication unit 116 identifies speaker US. The authentication unit 116 also obtains confidence information output from the confidence calculation unit 115. Based on the speaker information of the identified speaker US and the confidence information, the authentication unit 116 generates an authentication result screen SC and sends it to monitor MN.
[0036] Memory 12 includes, for example, Random Access Memory (hereinafter referred to as "RAM"), which serves as work memory used when executing various processes of the processor 11, and Read Only Memory (hereinafter referred to as "ROM"), which stores programs and data that define the operation of the processor 11. Data or information generated or acquired by the processor 11 is temporarily stored in RAM. Programs that define the operation of the processor 11 are written to ROM.
[0037] The feature extraction model database DB1 is a so-called storage system, configured using storage media such as flash memory, hard disk drive (hereinafter referred to as "HDD"), or solid state drive (hereinafter referred to as "SSD"). The feature extraction model database DB1 detects the speech segments of speaker US from registered voice data or authenticated voice data and stores a feature extraction model capable of extracting the features of this speaker US. The feature extraction model is a learned model generated by learning using, for example, deep learning.
[0038] The registered speaker database DB2 is a so-called storage system, configured using storage media such as flash memory, HDD, or SSD. The registered speaker database DB2 stores the feature quantities of multiple registered speakers that have been registered in advance, the results of the sound pickup condition determination corresponding to these feature quantities, and the registered speaker information in an associated manner.
[0039] The similarity calculation model database DB3 is a so-called storage system, configured using storage media such as flash memory, HDD, or SSD. The similarity calculation model database DB3 stores similarity calculation models capable of calculating the similarity between features extracted from authenticated speech data and features of registered speakers registered in the registered speaker database DB2. The similarity calculation models are trained and generated using training data under predetermined sound acquisition conditions through deep learning or the like.
[0040] For example, a similarity calculation model pre-trains and stores dimensions that tend to reveal individual characteristics in order to accurately calculate the similarity between two multidimensional vectors. It should be noted that the method of calculating similarity using a model is merely one example of a method for calculating similarity between vectors; existing techniques such as Euclidean distance and cosine similarity may also be used.
[0041] The sound acquisition condition-specific learning database DB4 is a so-called storage device, constructed using storage media such as flash memory, HDD, or SSD. The sound acquisition condition-specific learning database DB4 stores distribution information of the training data used to train the similarity calculation model.
[0042] The monitor MN is configured using a display such as a Liquid Crystal Display (LCD) or an organic electroluminescence (EL). The monitor MN displays the authentication result screen SC output from the terminal device P1.
[0043] The authentication result screen SC is a screen that notifies the administrator (for example, a person viewing monitor MN) of the speaker authentication result, and includes the authentication result information "Matches XX XX's voice." and the confidence level information "Confidence level: High". The authentication result screen SC may also include other registered speaker information (for example, a facial image). The authentication result screen SC does not have to include confidence level information.
[0044] Next, the operation procedure of terminal device P1 will be described with reference to Figure 3. Figure 3 is a flowchart showing an example of the operation procedure of terminal device P1 in the embodiment.
[0045] Terminal device P1 acquires audio data from microphone MK (St11). Microphone MK may be, for example, a microphone built into a PC, notebook PC, smartphone, or tablet device.
[0046] Terminal device P1 determines whether the control command associated with the voice data is a control command that requests registration in the registered speaker database DB2 (St12).
[0047] In step St12, terminal device P1 determines that if the control command is a control command requesting registration to the registered speaker database DB2, it will register the speaker US feature quantities in the registered speaker database DB2 (St12, YES), and extracts the speaker US feature quantities from the audio data (registered audio data) (St13).
[0048] On the other hand, in the processing of step St12, terminal device P1 determines that if the control command is a control command requesting speaker authentication rather than a control command requesting registration in the registered speaker database DB2, it will not register new speaker US feature quantities in the registered speaker database DB2 (St12, NO), and extracts speaker US feature quantities from the voice data (authentication voice data) (St14).
[0049] Terminal device P1 performs a process to determine the sound pickup conditions for the registered voice data (spoken voice) from which the speaker US features extracted from the registered voice data were extracted (St15A).
[0050] Terminal device P1 associates the speaker US feature quantities, sound acquisition condition information, and speaker US speaker information, and stores (registers) them in the registered speaker database DB2 (St16).
[0051] Terminal device P1 performs a determination process for the sound pickup conditions of the registered voice data (spoken voice) from which the speaker US features extracted from the authentication voice data have been extracted (St15B).
[0052] Terminal device P1 acquires information on the feature quantities and sound pickup conditions of multiple registered speakers registered in the registered speaker database DB2. Based on the information on the speaker US feature quantities and sound pickup conditions, and the information on the feature quantities and sound pickup conditions of multiple registered speakers, terminal device P1 selects a similarity calculation model to calculate the similarity between the speaker US feature quantities and the feature quantities of multiple registered speakers (St18).
[0053] Here, terminal device P1 refers to correspondence lists LST, LST1, and LST2 (see Figures 7, 8, and 9), which associate the selected similarity calculation model with the recording conditions of speaker US and the respective recording conditions of multiple registered speakers, and selects one of the multiple selection models.
[0054] Terminal device P1 performs speaker authentication processing based on the selected correspondence list LST, LST1, and LST2 selection models (St19).
[0055] Next, referring to Figures 4 and 5, the procedure for determining the sound pickup conditions shown in Step St15A and Step St15B in Figure 3 will be explained. Figure 4 is a flowchart showing an example of the procedure for determining the sound pickup conditions of the terminal device P1 in the embodiment. Figure 5 is a diagram illustrating an example of determining the sound pickup conditions and an example of calculating the reliability.
[0056] Terminal device P1 acquires speaker US feature quantities from the authenticated voice data (St151) and retrieves information on the training data for each similarity calculation model registered in the similarity calculation model database DB3 (St152). The training data referred to here is data used to calculate the similarity of feature quantities under predetermined sound acquisition conditions.
[0057] Terminal device P1 calculates the distance between, for example, the speaker US feature and each training data in order to select the optimal model (St153). Terminal device P1 identifies and selects the similarity calculation model with the smallest calculated distance (St154). Terminal device P1 determines that the sound pickup conditions corresponding to this similarity calculation model are the sound pickup conditions for the speaker US feature (St154).
[0058] For example, as shown in Figure 5, terminal device P1 measures the distance d between the speaker US feature quantity PT1 and the distribution of training data for the similarity calculation models DBA, DBB, and DBC. XA d XB d XC Calculate.
[0059] Note that while Figure 5 shows examples of the three similarity calculation models DBA, DBB, and DBC, each containing five training data points, the number of training data points used to generate (train) the similarity calculation model only needs to be one or more. Furthermore, the number of similarity calculation models DBA, DBB, and DBC only needs to be two or more.
[0060] Terminal device P1 uses the following (Equation 1) to calculate the distance d between the speaker US feature quantity PT1 and the similarity calculation models DBA, DBB, and DBC, respectively. XA d XB d XC The following is calculated. Note that the number N (N: an integer greater than or equal to 1) shown in (Equation 1) represents the number of training data used to generate the similarity calculation model, or the number of representative training data used to calculate distance among the similarity calculation models. Furthermore, the number of training data N for each similarity calculation model does not have to be the same.
[0061]
number
[0062] For example, in the example shown in Figure 5, N=5 and the distance d XAis the average distance between the feature quantity PT1 of the speaker US and the five learning data A1, A2, A3, A4, A5 included in the similarity calculation model DBA. The distance d XB is the average distance between the feature quantity PT1 of the speaker US and the five learning data B1, B2, B3, B4, B5 included in the similarity calculation model DBB. The distance d XC is the average distance between the feature quantity PT1 of the speaker US and the five learning data C1, C2, C3, C4, C5 included in the similarity calculation model DBC.
[0063] The terminal device P1 selects the similarity calculation model corresponding to the distribution of the learning data, which is the minimum distance among the calculated multiple distances, as the similarity calculation model to be used for calculating the similarity. For example, in the example shown in FIG. 5, the calculated distances d XA , d XB , d XC among which the distance d XC is the minimum distance. In such a case, the terminal device P1 determines that the sound collection condition of the uttered voice (feature quantity PT1) of the speaker US is the same as the sound collection condition corresponding to the similarity calculation model DBC.
[0064] Further, the terminal device P1 calculates the reliability of the calculated similarity based on the distance between the distribution of the learning data of the similarity calculation model used for calculating the similarity and the feature quantity of the speaker US (that is, the distances d XA , d XB , d XC ).
[0065] The terminal device P1 determines whether the distance between the similarity calculation model used for calculating the similarity and the feature quantity of the speaker US is less than or equal to a predetermined value.
[0066] When the terminal device P1 determines that the distance between the distribution of the learning data of the similarity calculation model and the feature quantity of the speaker US is less than or equal to a predetermined value, it calculates (evaluates) the reliability of the similarity as "high". On the other hand, when the terminal device P1 determines that the distance between the distribution of the learning data of the similarity calculation model and the feature quantity of the speaker US is not less than or equal to a predetermined value, it calculates (evaluates) the reliability of the similarity as "low".
[0067] For example, in the example shown in Figure 5, terminal device P1 calculates the distance d between the similarity calculation model DBC and the speaker US feature quantity PT1. XC If the distance d is below a predetermined value, the confidence level of the similarity is calculated (evaluated) as "high". Meanwhile, terminal device P1 calculates the distance d between the similarity calculation model DBC and the speaker US feature quantity PT1. XC If the value is not below a predetermined value, the confidence level of the similarity is calculated (evaluated) as "low".
[0068] Next, with reference to Figure 6, the speaker authentication procedure shown in step St19 in Figure 3 will be described. Figure 6 is a flowchart showing an example of the speaker authentication procedure for terminal device P1 in the embodiment.
[0069] Terminal device P1 reads a similarity calculation model from the similarity calculation model database DB3, which is used to determine the similarity between the speaker US feature quantity and the feature quantity of any of the registered speakers (St191), based on the correspondence lists LST, LST1, and LST2.
[0070] Terminal device P1 calculates the similarity between the speaker US feature vectors and the registered speaker feature vectors using a similarity calculation model (St192). Terminal device P1 also calculates the confidence level of the calculated similarity based on the distance between the similarity calculation model calculated using (Equation 1) and the speaker US feature vectors (St192). Terminal device P1 repeatedly executes the process in step St192 until it has calculated the similarity and confidence levels between the speaker US speech data feature vectors and the feature vectors of all registered speakers registered in the registered speaker database DB2.
[0071] Terminal device P1 determines whether any of the calculated similarities are equal to or above a threshold (St193).
[0072] In step St193, if terminal device P1 determines that any of the calculated similarities are above a threshold (St193, YES), it determines that the registered speaker and speaker US corresponding to the similarity determined to be above the threshold are the same person. Based on the registered speaker information of this registered speaker, terminal device P1 identifies speaker US (St194). If there are multiple similarities determined to be above the threshold, terminal device P1 may determine that the registered speaker and speaker US corresponding to the highest calculated similarity are the same person.
[0073] If terminal device P1 determines in step St193 that none of the calculated similarities are equal to or above a threshold (St193, NO), it determines that speaker US cannot be identified (St195).
[0074] Terminal device P1 generates an authentication result screen SC based on the registered speaker information of the identified speaker US. Terminal device P1 outputs the generated authentication result screen SC to monitor MN for display (St196).
[0075] As described above, terminal device P1 registers speaker information, speaker US feature quantities, and speaker US feature quantity sound pickup conditions in association with each other during voice registration. By registering sound pickup conditions that change the feature quantities that indicate the individuality of speaker US, terminal device P1 can perform speaker authentication with higher accuracy even if the sound pickup conditions for the feature quantities at voice registration differ from those at voice authentication, and even if the feature quantities at voice registration and voice authentication differ, by selecting a similarity calculation model based on the sound pickup conditions for each feature quantity. Therefore, terminal device P1 can more effectively suppress the decrease in speaker authentication accuracy caused by noise contained in the authenticated voice data.
[0076] Furthermore, terminal device P1 calculates and displays a confidence level indicating the likelihood of the speaker identified through the speaker authentication process, based on the similarity calculation model used to calculate the similarity. This allows terminal device P1 to present the likelihood of the speaker authentication result to the administrator viewing monitor MN. Therefore, terminal device P1 can inform the administrator, by displaying the confidence level, that there was no similarity calculation model suitable for calculating the similarity, and that speaker authentication was performed using the "general-purpose model" similarity calculation model described later.
[0077] Examples of correspondence lists LST1 and LST2 will be explained with reference to Figures 7 and 8, respectively. Figure 7 is a diagram illustrating an example of correspondence list LST1 when the estimated sound pickup conditions during voice registration and the estimated sound pickup conditions during voice authentication are the same. Figure 8 is a diagram illustrating an example of correspondence list LST2 when the sound pickup conditions during voice registration and the sound pickup conditions during voice authentication are different.
[0078] Figures 7 and 8 illustrate an example of referencing correspondence lists LST1 and LST2 based on the information of the sound pickup conditions associated with the feature quantities of one registered speaker registered in the registered speaker database DB2, and the sound pickup conditions of the authenticated speech data of speaker US, who is the target of speaker authentication, in order to make the explanation easier to understand.
[0079] In the example of the correspondence list LST1 shown in Figure 7, the voice data used during voice registration and voice authentication are the same under the same sound acquisition conditions: "in-store background noise, male voice, telephone voice".
[0080] Terminal device P1 extracts speaker US features from the authenticated voice data of speaker US transmitted from a sound pickup device such as microphone MK. Terminal device P1 refers to the learning database DB4 for each sound pickup condition and determines that the sound pickup condition for speaker US features is "store noise, male, telephone voice" based on the extracted features and the distribution of learning data for each sound pickup condition.
[0081] Furthermore, terminal device P1 selects a similarity calculation model corresponding to the combination of speaker US's sound pickup conditions and the sound pickup conditions of multiple registered speakers registered in the registered speaker database DB2. Terminal device P1 refers to correspondence list LST1, which associates speaker US's sound pickup conditions, the sound pickup conditions of multiple registered speakers, and the selected similarity calculation model.
[0082] The correspondence list LST1 is data that associates the "Sound Pickup Condition Judgment Result 1," which is the result of determining the sound pickup conditions for speaker US features, with the "Sound Pickup Condition Judgment Result 2," which is the result of determining the sound pickup conditions for registered speakers registered in the registered speaker database DB2, and the "Selected Model," which is a similarity calculation model selected based on these two sound pickup conditions.
[0083] The "Sound Acquisition Condition Judgment Result 1" shows the sound acquisition condition judgment result for speaker US.
[0084] The "Sound Collection Condition Judgment Result 2" shows the judgment result for the sound collection conditions of registered speakers registered in the registered speaker database DB2. Information on the judgment probability corresponding to each sound collection condition judgment result is not mandatory and may be omitted.
[0085] Furthermore, the results of the sound pickup condition determination may include multiple conditions, such as the gender of the speaker (US), the sound pickup device, and the type of noise.
[0086] The similarity calculation model "Selected Model" includes the similarity calculation models "Model A," "Model B," "Model C," and "Model Z," which are selected in accordance with the sound collection condition determination result "Sound Collection Condition Determination Result 1" and the sound collection condition determination result "Sound Collection Condition Determination Result 2."
[0087] For example, the similarity calculation model "Model A" is a similarity calculation model that has been determined to be optimal for calculating the similarity between the speaker US features and the registered speaker features when the sound pickup condition for the speaker US features is "XX1" and the information for the sound pickup condition for the registered speaker features is "XX1".
[0088] Furthermore, if the similarity calculation model selection unit 114 determines that there is no similarity calculation model suitable for the similarity calculation process based on the combination of the sound acquisition conditions for speaker US features and the sound acquisition conditions for registered speaker features, it selects the general-purpose similarity calculation model "Model Z".
[0089] The similarity calculation model selection unit 114 selects a similarity calculation model database from the similarity calculation model database DB3 based on the referenced correspondence list LST1. For example, in the example shown in Figure 7, the similarity calculation model selection unit 114 selects the similarity calculation model "Model A".
[0090] Next, in the example of the correspondence list LST2 shown in Figure 8, the registered voice data during voice registration has the same sound pickup conditions: "in-store noise, male, telephone voice". The voice data during voice authentication has different sound pickup conditions from the registered voice data during voice registration: "outdoor noise, male, headset voice".
[0091] Terminal device P1 extracts speaker US features from the authenticated voice data of speaker US transmitted from a sound pickup device such as microphone MK. Terminal device P1 refers to the learning database DB4 for each sound pickup condition and determines that the sound pickup condition for speaker US features is "store noise, male, telephone voice" based on the extracted features and the distribution of learning data for each sound pickup condition.
[0092] Furthermore, terminal device P1 selects a similarity calculation model corresponding to the combination of speaker US's sound pickup conditions and the sound pickup conditions of multiple registered speakers registered in the registered speaker database DB2. Terminal device P1 refers to correspondence list LST2, which associates speaker US's sound pickup conditions, the sound pickup conditions of multiple registered speakers, and the selected similarity calculation model.
[0093] The correspondence list LST2 is data that associates the "Sound Acquisition Condition Judgment Result 3," which is the result of determining the sound acquisition conditions for speaker US features, the "Sound Acquisition Condition Judgment Result 4," which is the result of determining the sound acquisition conditions for registered speakers registered in the registered speaker database DB2, and the "Selected Model," which is a similarity calculation model selected based on these two sound acquisition conditions.
[0094] The "Sound Acquisition Condition Judgment Result 3" shows the sound acquisition condition judgment result for speaker US.
[0095] The "Sound Collection Condition Judgment Result 4" shows the judgment result for the sound collection conditions of registered speakers registered in the registered speaker database DB2. Information on the judgment probability corresponding to each sound collection condition judgment result is not mandatory and may be omitted.
[0096] Furthermore, the results of the sound pickup condition determination may include multiple conditions, such as the gender of the speaker (US), the sound pickup device, and the type of noise.
[0097] The similarity calculation model "Selected Model" includes the similarity calculation models "Model D," "Model E," "Model F," and "Model Z," respectively, which are selected in accordance with the sound collection condition determination result "Sound Collection Condition Determination Result 3" and the sound collection condition determination result "Sound Collection Condition Determination Result 4."
[0098] For example, the similarity calculation model "Model D" is a similarity calculation model that was determined to be optimal for calculating the similarity between the speaker US features and the registered speaker features when the sound pickup conditions for the speaker US features are "XX1" and the information for the sound pickup conditions for the registered speaker features is "XX4".
[0099] Furthermore, the similarity calculation model selection unit 114, based on the combination of the speaker US feature acquisition conditions and the registered speaker feature acquisition conditions, selects the "general-purpose model" if it determines that there is no similarity calculation model suitable for the similarity calculation process.
[0100] The similarity calculation model selection unit 114 selects a similarity calculation model database from the similarity calculation model database DB3 based on the referenced correspondence list LST2. For example, in the example shown in Figure 8, the similarity calculation model selection unit 114 selects the similarity calculation model "Model D".
[0101] As described above, terminal device P1 can select the most suitable similarity calculation model for calculating the similarity of two features (speaker US features and registered speaker features) based on the combination of sound pickup conditions included in the speaker US features and registered speaker features, respectively, which are the targets for similarity calculation. This allows terminal device P1 to select the most suitable similarity calculation model for calculating the similarity of the two features even if the sound pickup conditions included in the features at the time of voice registration and the features at the time of voice authentication change. In other words, terminal device P1 can more effectively suppress the decrease in speaker authentication accuracy caused by noise in the voice data. When there are multiple candidate conditions for sound pickup condition determination, terminal device P1 may, for example, calculate the similarity using the corresponding similarity calculation model for each and then calculate the average value to adopt as the similarity.
[0102] Refer to Figure 9 for a detailed explanation of the correspondence list. Figure 9 is a diagram illustrating a specific example of the correspondence list (LST).
[0103] For clarity, Figure 9 explains an example where the voice data for voice registration and voice authentication are the same under the same sound pickup conditions: "store background noise, male voice, telephone voice," similar to the example of the correspondence list LST1 shown in Figure 7.
[0104] The correspondence list LST associates the "Acquisition Condition Judgment Result AA" which is the result of the speaker US feature quantity sound acquisition condition judgment, the "Acquisition Condition Judgment Result BB" which is information on the registered speaker feature quantity sound acquisition condition, and the selected similarity calculation model "Selected Model".
[0105] The "Sound Acquisition Condition Judgment Result AA" is the result of determining the sound acquisition conditions for the speaker's US features, and the "Sound Acquisition Condition Judgment Result BB" is the information on the sound acquisition conditions for the registered speaker's features. For example, the sound acquisition conditions include "gender," "equipment (sound quality)," and "noise type." Note that the sound acquisition conditions shown in Figure 9 are just examples and are not limited to them. In addition, in the correspondence list LST, the same sound acquisition conditions are shown in bold in both the "Sound Acquisition Condition Judgment Result AA" and the "Sound Acquisition Condition Judgment Result BB" information.
[0106] The "Gender" setting in the sound acquisition conditions indicates the gender of the speaker US, determined based on the speaker US features. The "Equipment (Sound Quality)" setting in the sound acquisition conditions indicates information about the sound acquisition device, determined based on the speaker US features. The "Noise Type" setting in the sound acquisition conditions indicates the type of ambient sound, noise, etc., included in the speaker US features during speech.
[0107] The "Selection Model" similarity calculation model is used to calculate the similarity between the features of Speaker US and the features of registered speakers.
[0108] For example, terminal device P1 may determine that the result of two sound pickup conditions is that the "gender" condition is "male", the "device (sound quality)" condition is "telephone", and the "noise type" condition is "in-store noise". In such a case, terminal device P1 selects a similarity calculation model, "Male Telephone In-store Noise Model," which is suitable for calculating the similarity of the same sound pickup conditions "male", "telephone", and "in-store noise" features in the two sound pickup conditions.
[0109] For example, terminal device P1 may determine that, for instance, the results of two sound pickup conditions are: the "gender" sound pickup condition is "male," the "device (sound quality)" sound pickup condition is "telephone" or "headset," and the "noise type" sound pickup condition is "in-store noise" or "outdoor noise." In such a case, terminal device P1 selects a similarity calculation model, the "male model," which is suitable for calculating the similarity of feature quantities for the "male" sound pickup condition that are identical or similar to the two sound pickup conditions.
[0110] For example, terminal device P1 may determine that, for instance, the results of two sound pickup conditions are: "Gender" is "Female", "Equipment (Sound Quality)" is "Headset", and "Noise Type" is "None (Clean Voice)" and "Store Noise". In such a case, terminal device P1 selects a similarity calculation model, "Female Headset Model," which is suitable for calculating the similarity of the feature quantities of the identical or similar sound pickup conditions "Female" and "Headset".
[0111] As described above, the terminal device P1 according to the embodiment includes a communication unit 10 (an example of an acquisition unit) that acquires voice data, a feature extraction unit 111 (an example of a detection unit) that detects the utterance section spoken by speaker US from the voice data, a feature extraction unit 111 (an example of an extraction unit) that extracts speaker US features (an example of utterance features) from the detected utterance section, a similarity calculation model selection unit 114 (an example of a selection unit) that selects a first similarity calculation model (an example of a first similarity calculation model, which is selected in the process of step St18) from among a plurality of similarity calculation models (an example of a similarity calculation model) to be used for authenticating speaker US, based on the extracted speaker US features and the features of at least one registered speaker that has been registered in advance, and an authentication unit 116 that authenticates speaker US by comparing the speaker US features with the registered speaker features using the selected first similarity calculation model.
[0112] As a result, the terminal device P1 according to the embodiment can select a similarity calculation model suitable for the similarity calculation process of two feature quantities (speaker US feature quantity and registered speaker feature quantity), and by using the selected similarity calculation model, it can perform speaker authentication based on the two feature quantities with higher accuracy. In other words, the terminal device P1 can more effectively suppress the decrease in speaker authentication accuracy caused by the difference in noise (environmental noise) between voice registration and voice authentication.
[0113] Furthermore, the terminal device P1 according to the embodiment further includes a sound pickup condition determination unit 112 (an example of a determination unit) that determines the sound pickup conditions for spoken speech corresponding to the feature quantities based on the feature quantities. The sound pickup condition determination unit 112 selects a first similarity calculation model based on the sound pickup conditions corresponding to the feature quantities of speaker US and the sound pickup conditions corresponding to the feature quantities of the registered speaker. As a result, the terminal device P1 according to the embodiment can select a similarity calculation model that is more suitable for matching the feature quantities based on the combination of the sound pickup conditions at the time of speech registration and the sound pickup conditions at the time of speech authentication. Therefore, the terminal device P1 can more effectively suppress the decrease in speaker authentication accuracy.
[0114] Furthermore, the multiple similarity calculation models in the terminal device P1 according to the embodiment are generated using at least one training data under predetermined sound pickup conditions. The sound pickup condition determination unit 112 determines the sound pickup conditions of the speaker US or registered speaker based on the distance between the feature quantities of the speaker US or registered speaker and each of the multiple similarity calculation models. As a result, the terminal device P1 according to the embodiment can determine the sound pickup conditions corresponding to each feature quantity with higher accuracy based on the distance between the similarity calculation model generated using the training data for each sound pickup condition and the feature quantities of the speaker US or registered speaker.
[0115] Furthermore, the sound pickup condition determination unit 112 in the terminal device P1 according to the embodiment selects a second similarity calculation model (an example of a second similarity calculation model, selected in the process of step St154) which has the shortest distance between the feature quantities of speaker US or registered speaker and each of the multiple similarity calculation models, and determines the sound pickup conditions of speaker US or registered speaker based on the sound pickup conditions corresponding to the selected second similarity calculation model. As a result, the terminal device P1 according to the embodiment can determine the sound pickup conditions corresponding to the feature quantities of speaker US or registered speaker with higher accuracy because it selects the similarity calculation model with the closest characteristics when calculating the similarity with the feature quantities of speaker US or registered speaker.
[0116] Furthermore, the terminal device P1 according to the embodiment further includes an authentication unit 116 (an example of a calculation unit) that calculates the similarity between the feature quantities of the speech data of the utterance section and the respective feature quantities of multiple registered speakers. The authentication unit 116 authenticates speaker US based on the multiple similarity values calculated. As a result, the terminal device P1 according to the embodiment can perform speaker authentication using the similarity between the feature quantities of multiple registered speakers registered in advance and the feature quantities of speaker US.
[0117] Furthermore, the authentication unit 116 in the terminal device P1 according to the embodiment identifies a registered speaker whose similarity is above a threshold as speaker US. As a result, the terminal device P1 according to the embodiment can perform speaker authentication using the similarity between the feature quantities of multiple registered speakers registered in advance and the feature quantities of speaker US.
[0118] Furthermore, the authentication unit 116 in the terminal device P1 according to the embodiment generates and outputs an authentication result screen SC that includes information about registered speakers whose similarity is above a threshold. This allows the terminal device P1 according to the embodiment to present the speaker authentication result to the speaker US or the administrator.
[0119] Furthermore, if the authentication unit 116 in the terminal device P1 according to the embodiment determines that the calculated similarity scores are not above a threshold, it determines that speaker US cannot be identified. As a result, the terminal device P1 according to the embodiment can more effectively suppress a decrease in speaker authentication accuracy and more effectively suppress misidentification of speaker US.
[0120] Furthermore, the terminal device P1 according to the embodiment further includes an authentication unit 116 that calculates the similarity between the feature quantities of the speech data of the utterance section and the respective feature quantities of multiple registered speakers, and a reliability calculation unit 115 (an example of a reliability calculation unit) that calculates the reliability of the similarity. The reliability calculation unit 115 calculates the reliability of the similarity based on the distance between the speaker US feature quantities and the second similarity calculation model. As a result, the terminal device P1 according to the embodiment can calculate the reliability of the similarity, or the sound pickup conditions determined in the process of calculating the similarity, or the first similarity calculation model used to calculate the similarity.
[0121] Furthermore, the authentication unit 116 in the terminal device P1 according to the embodiment identifies registered speakers whose similarity is above a threshold as speaker US, and generates and outputs an authentication result screen SC that includes information about registered speakers whose similarity is above a threshold and information about the calculated confidence level. As a result, the terminal device P1 according to the embodiment can prompt the administrator to confirm whether or not the speaker authentication result is reliable by displaying the speaker authentication result and the confidence level of the speaker authentication result.
[0122] Furthermore, in the terminal device P1 according to the embodiment, the sound pickup conditions include at least one of the speaker US's gender, the speaker US's age, the speaker US's language, the sound pickup device used to capture the spoken voice, or the type of noise contained in the spoken voice. This allows the terminal device P1 according to the embodiment to change the feature quantities used for speaker authentication and select a similarity calculation model based on the sound pickup conditions that may cause a decrease in speaker authentication accuracy. Therefore, the terminal device P1 can more effectively suppress the decrease in speaker authentication accuracy.
[0123] Although various embodiments have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the various embodiments described above can be combined arbitrarily without departing from the spirit of the invention.
[0124] This application is based on a Japanese patent application (Patent Application No. 2022-045391) filed on March 22, 2022, the contents of which are incorporated by reference within this application. [Industrial applicability]
[0125] This disclosure is useful as a voice authentication device and voice authentication method that can suppress the decline in speaker authentication accuracy caused by changes in ambient noise. [Explanation of Symbols]
[0126] 10 Communications Department 11 processors 12 memory 100 Voice Recognition Systems 111 Feature Extraction Unit 112 Sound collection condition determination section 113 Speaker Registration Department 114 Similarity Calculation Model Selection Section 115 Reliability Calculation Unit 116 Authentication Department DB1 Feature Extraction Model Database DB2 Registered Speaker Database DB3 Similarity Calculation Model Database DB4 Learning Database Based on Sound Collection Conditions MK Microphone MN Monitor P1 Terminal Device SC Authentication Results Screen US speaker
Claims
1. An acquisition unit that acquires audio data, A detection unit that detects the speech segment spoken by the speaker from the aforementioned audio data, An extraction unit that extracts the speaker's speech features from the detected speech segment, A determination unit that determines the sound pickup conditions for the speech corresponding to the speech features based on the speech features, A selection unit selects a first similarity calculation model from among multiple similarity calculation models to be used for speaker authentication, based on the extracted speech features of the speaker and the speech features of multiple registered speakers that have been registered in advance. The system includes an authentication unit that authenticates the speaker by comparing the speech features of the speaker with the speech features of the registered speaker using a selected first similarity calculation model, The selection unit selects the first similarity calculation model based on the sound acquisition conditions corresponding to the speech features of the speaker and the sound acquisition conditions corresponding to the speech features of the registered speaker. Voice authentication device.
2. The aforementioned multiple similarity calculation models are generated using at least one training data under predetermined sound acquisition conditions. The determination unit determines the sound pickup conditions of the speaker or registered speaker based on the average distance between the speech features of the speaker or registered speaker and the multiple training data used to train one of the multiple similarity calculation models. The voice authentication device according to claim 1.
3. The determination unit calculates the average distance between the speech features of the speaker or registered speaker and the multiple training data used to train each of the multiple similarity calculation models, selects the second similarity calculation model with the shortest average distance, and determines the sound pickup conditions of the speaker or registered speaker based on the sound pickup conditions corresponding to the selected second similarity calculation model. The voice authentication device according to claim 2.
4. The system further comprises a calculation unit that calculates the similarity between the speech feature quantities of the speech data in the speech interval and the respective speech feature quantities of the multiple registered speakers that have been registered in advance. The authentication unit authenticates the speaker based on the calculated multiple similarity scores. The voice authentication device according to claim 1.
5. The authentication unit identifies a registered speaker whose similarity is equal to or greater than a threshold as the speaker. The voice authentication device according to claim 4.
6. The authentication unit generates and outputs an authentication result screen that includes information about the registered speaker whose similarity is equal to or greater than the threshold. The voice authentication device according to claim 5.
7. If the authentication unit determines that the calculated similarity scores are not equal to or greater than a threshold, it determines that the speaker cannot be identified. The voice authentication device according to claim 5.
8. A calculation unit that calculates the similarity between the speech feature quantities of the speech data in the speech interval and the respective speech feature quantities of the multiple registered speakers that have been registered in advance, The system further comprises a confidence calculation unit for calculating the confidence level of the similarity, The confidence calculation unit calculates the confidence level of the similarity based on the average distance between the speaker's speech features and the plurality of training data used to train the second similarity calculation model. The voice authentication device according to claim 3.
9. The authentication unit identifies registered speakers whose similarity is equal to or greater than a threshold as the speaker, and generates and outputs an authentication result screen that includes information about the registered speakers whose similarity is equal to or greater than the threshold, and the calculated confidence level information. The voice authentication device according to claim 8.
10. The sound recording conditions include at least one of the speaker's gender, the speaker's age, the speaker's language, the sound recording device used to record the speech, or the type of noise contained in the speech. The voice authentication device according to claim 1.
11. A voice authentication method performed by a terminal device, Acquire audio data, From the aforementioned audio data, the speech segment spoken by the speaker is detected, The speech features of the speaker are extracted from the detected speech segments. Based on the aforementioned speech features, the conditions for capturing the speech corresponding to the aforementioned speech features are determined. Based on the sound pickup conditions corresponding to the extracted speech features of the speaker and the sound pickup conditions corresponding to the speech features of multiple registered speakers registered in advance, a first similarity calculation model is selected from among multiple similarity calculation models to be used for speaker authentication. Using the selected first similarity calculation model, the speaker's utterance features are compared with the registered speaker's utterance features to authenticate the speaker. Voice authentication method.