Voice authentication device and voice authentication method

The voice authentication device and method enhance speaker authentication accuracy by identifying speech and non-speech sections and selecting appropriate similarity calculation models based on noise types, addressing the issue of environmental noise variability.

JP7742568B2Active Publication Date: 2025-09-22PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024510013
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-22
Filing Date
2023-03-10
Publication Date
2025-09-22
Estimated Expiration
2043-03-10

AI Technical Summary

Technical Problem

Existing voiceprint authentication systems suffer from decreased accuracy due to variations in environmental noise, as the feature quantities extracted from voice signals can vary based on the noise present, leading to misidentification of individuals.

Method used

A voice authentication device and method that includes an acquisition unit, detection unit, extraction unit, selection unit, and authentication unit to identify speech and non-speech sections, select appropriate similarity calculation models based on registered noise features, and compare speech features to authenticate speakers, thereby mitigating the impact of environmental noise.

Benefits of technology

The solution effectively suppresses the decrease in speaker authentication accuracy caused by changes in environmental noise, ensuring more accurate identification by selecting optimal similarity calculation models based on noise types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742568000001
    Figure 0007742568000001
  • Figure 0007742568000002
    Figure 0007742568000002
  • Figure 0007742568000003
    Figure 0007742568000003
Patent Text Reader

Abstract

This voice authentication device comprises: a detection unit for detecting, from speech data, a speech segment in which a speaker is speaking and a non-speech segment in which the speaker is not speaking; an extraction unit for extracting a speech feature amount of the speech segment and noise contained in the non-speech segment; a selection unit for selecting, on the basis of the extracted noise and noise associated with a plurality of pre-registered registered feature amounts, one similarity calculation model from among a plurality of similarity calculation models; and an authentication unit for authenticating the speaker by matching the speech feature amounts of the speaker with the registered feature amounts of the registered speaker using the selected similarity calculation model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a voice authentication device and a voice authentication method. [Background technology]

[0002] Patent Document 1 discloses a speech recognition device that recognizes the speech of a subject. The speech recognition device stores a plurality of action noise models created corresponding to each of a plurality of actions, each associated with the plurality of actions, detects input speech including the subject's speech, identifies the subject's action, and reads out the action noise model corresponding to the action identified by the action identification means. The speech recognition device reads out an environmental noise model corresponding to the subject's current position, combines the read-out action noise model with the environmental noise model, and recognizes the subject's speech included in the detected input speech using the combined noise superposition model. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2008-250059 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in Patent Document 1, it was necessary to collect in advance the motion noises generated by each of a plurality of motions and the environmental noises at each of a plurality of positions where voice recognition can be performed, which was very time-consuming. Also, in voiceprint authentication, the feature quantity indicating the individuality of the person to be authenticated, extracted from the voice signal, varies depending on the noise contained in the voice signal. Therefore, when performing voiceprint authentication using the above-mentioned voice recognition device, if the noise contained in the pre-registered voice signal and the noise contained in the voice signal collected during voiceprint authentication are different, the feature quantities extracted from the respective voice signals may not indicate the individuality of the same person, which may result in a decrease in the accuracy of voiceprint authentication.

[0005] The present disclosure has been devised in view of the above-described conventional situation, and aims to provide a voice authentication device and a voice authentication method that suppress a decrease in speaker authentication accuracy caused by changes in environmental noise. [Means for solving the problem]

[0006] The present disclosure provides a voice authentication device including: an acquisition unit that acquires voice data; a detection unit that detects, from the voice data, speech sections in which a speaker is speaking and non-speech sections in which the speaker is not speaking; an extraction unit that extracts speech features of the speech sections and noise included in the non-speech sections of the voice data; a selection unit that selects one of a plurality of similarity calculation models based on the extracted noise and noise associated with registered features of a plurality of registered speakers that have been registered in advance; and an authentication unit that uses the selected similarity calculation model to compare the speech features of the speaker with the registered features of the registered speakers to authenticate the speaker.

[0007] The present disclosure also provides a voice authentication method performed by a terminal device, which acquires voice data, detects from the voice data speech sections in which a speaker is speaking and non-speech sections in which the speaker is not speaking, extracts speech features of the speech sections and noise included in the non-speech sections of the voice data, selects one of similarity calculation models based on the extracted noise and noise associated with registered features of a plurality of registered speakers who have been registered in advance, and uses the selected similarity calculation model to compare the speech features of the speaker with the registered features of the registered speakers to authenticate the speaker. [Effects of the Invention]

[0008] According to the present disclosure, it is possible to suppress a decrease in speaker authentication accuracy caused by changes in environmental noise. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a block diagram showing an example of the internal configuration of a voice authentication system according to an embodiment. [Figure 2] FIG. 1 is a diagram illustrating each process performed by a processor of a terminal device according to an embodiment. [Figure 3] 1 is a flowchart showing an example of an operation procedure of a terminal device according to an embodiment; [Figure 4] 1 is a flowchart showing an example of a speaker authentication procedure of a terminal device according to an embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of a correspondence list when the noise type at the time of voice registration is the same as the noise type at the time of voice authentication. [Figure 6] FIG. 10 is a diagram illustrating an example of a correspondence list when the noise type at the time of voice registration differs from the noise type at the time of voice authentication. [Figure 7] A diagram illustrating an example of reliability calculation. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, with reference to the accompanying drawings as appropriate, detailed descriptions of embodiments specifically disclosing a voice authentication device and a voice authentication method according to the present disclosure will be provided. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter recited in the claims.

[0011] (Embodiment) First, a voice authentication system 100 according to an embodiment will be described with reference to Fig. 1 and Fig. 2. Fig. 1 is a block diagram showing an example of the internal configuration of the voice authentication system 100 according to the embodiment. Fig. 2 is a diagram explaining each process performed by the processor 11 of the terminal device P1 according to the embodiment.

[0012] The voice authentication system 100 includes a terminal device P1 as an example of a voice authentication device, a monitor MN, a noise determination device P2, and a network NW. The voice authentication system 100 may also include a microphone MK or a monitor MN.

[0013] The microphone MK collects the speech of the speaker US for pre-registration in the terminal device P1. The microphone MK converts the collected speech of the speaker US into a voice signal or voice data to be registered in the terminal device P1. The microphone MK transmits the converted voice signal or voice data to the processor 11 via the communication unit 10.

[0014] The microphone MK also collects the speech of the speaker US, which is used for speaker authentication. The microphone MK converts the collected speech of the speaker US into a voice signal or voice data. The microphone MK transmits the converted voice signal or voice data to the processor 11 via the communication unit 10.

[0015] In the following explanation, for ease of understanding, voice data for voice registration, or voice data already registered in terminal device P1, will be referred to as "registration voice data," and voice data for voice authentication will be referred to as "authentication voice data."

[0016] The microphone MK may be a microphone provided in a predetermined device such as a personal computer (hereinafter referred to as "PC"), a notebook PC, a smartphone, a tablet terminal, etc. The microphone MK may also transmit an audio signal or audio data to the terminal device P1 by wireless communication via a network (not shown).

[0017] The terminal device P1 is realized by, for example, a PC, a notebook PC, a smartphone, a tablet terminal, etc., and executes a voice registration process using registration voice data of a speaker US and a speaker authentication process using authentication voice data. The terminal device P1 includes a communication unit 10, a processor 11, a memory 12, a feature extraction model database DB1, a registration speaker database DB2, a similarity calculation model database DB3, and a noise-similarity calculation model correspondence list DB4.

[0018] The communication unit 10, which is an example of an acquisition unit, is connected to the microphone MK, the monitor MN, and the noise determination device P2 so as to be able to transmit and receive data between them via wired or wireless communication. The wireless communication here refers to short-range wireless communication such as Bluetooth (registered trademark) or NFC (registered trademark), or communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark).

[0019] The communication unit 10 may transmit and receive data to and from the microphone MK via an interface such as Universal Serial Bus (USB), or may transmit and receive data to and from the monitor MN via an interface such as High-Definition Multimedia Interface (HDMI, registered trademark).

[0020] The processor 11 is configured using, for example, a central processing unit (CPU) or a field programmable gate array (FPGA), and performs various processes and controls in cooperation with the memory 12. Specifically, the processor 11 references the programs and data stored in the memory 12 and executes the programs to realize the functions of each unit, such as a noise extraction unit 111, a feature extraction unit 112, a noise determination unit 113, a speaker registration unit 114, a similarity calculation model selection unit 115, a reliability calculation unit 116, and an authentication unit 117.

[0021] When registering the voice of speaker US, processor 11 performs a new registration (storage) process of features of speaker US in registered speaker database DB2 by implementing the functions of noise extraction unit 111, feature extraction unit 112, noise determination unit 113, and speaker registration unit 114. Note that the features referred to here are features that indicate the individuality of speaker US, extracted from the registration voice data.

[0022] In addition, during voice authentication of speaker US, processor 11 executes speaker authentication processing by realizing the functions of each of the noise extraction unit 111, feature extraction unit 112, noise determination unit 113, similarity calculation model selection unit 115, reliability calculation unit 116, and authentication unit 117.

[0023] Noise extraction unit 111, which is an example of a detection unit and extraction unit, acquires registration voice data or authentication voice data of speaker US transmitted from microphone MK. Noise extraction unit 111 detects speech sections in which speaker US is speaking and sections in which speaker US is not speaking (hereinafter referred to as "non-speech sections") from the registration voice data or authentication voice data. Noise extraction unit 111 extracts noise contained in the detected non-speech sections and outputs data of the extracted noise (hereinafter referred to as "noise data") to noise determination unit 113.

[0024] Feature extraction unit 112, an example of a detection unit, acquires enrollment voice data or authentication voice data of speaker US transmitted from microphone MK. Feature extraction unit 112 detects speech segments from the enrollment voice data or authentication voice data, and extracts features indicative of the individuality of speaker US from the detected speech segments of the enrollment voice data or authentication voice data using a feature extraction model.

[0025] If the enrollment speech data of speaker US transmitted from microphone MK is associated with a control command requesting the enrollment of features of speaker US and the enrollment speaker information of speaker US, feature extraction unit 112 proceeds to the speech enrollment process of speaker US. Feature extraction unit 112 outputs the extracted features of speaker US to speaker enrollment unit 114.

[0026] Furthermore, when a control command requesting speaker authentication is associated with the authentication voice data of speaker US transmitted from microphone MK, feature extraction unit 112 executes speaker authentication processing. Feature extraction unit 112 outputs the extracted features of speaker US to authentication unit 117.

[0027] The noise determination unit 113 acquires the noise data output from the noise extraction unit 111. The noise determination unit 113 transmits the noise data to the noise determination device P2 via the network NW, and determines the type of noise contained in the enrollment voice data or authentication voice data of the speaker US.

[0028] Note that the noise referred to here is noise picked up due to the environment (background) at the time of sound collection, such as surrounding conversations, music, vehicle noise, wind noise, etc. The noise type indicates the environment (place), location, etc. where the noise occurs, such as noise inside a store, wind noise outside, music inside a store, inside a train station, etc. The noise type may further include information on the time of day, such as early morning, daytime, or nighttime.

[0029] In the voice registration process of speaker US, the noise determination unit 113 outputs noise type information corresponding to the noise type determination result transmitted from the noise determination device P2 to the speaker registration unit 114. In addition, in the speaker authentication process, the noise determination unit 113 outputs noise type information corresponding to the noise type determination result transmitted from the noise determination device P2 to the similarity calculation model selection unit 115.

[0030] The speaker registration unit 114 acquires the features of the speaker US output from the feature extraction unit 112 and the information on the noise type included in the enrollment voice data of the speaker US output from the noise determination unit 113. The speaker registration unit 114 associates the features of the speaker US, the information on the noise type, and the speaker information of the speaker US, and registers them in the enrollment speaker database DB2.

[0031] The speaker information may be extracted from registered voice data by voice recognition, or may be acquired from a device owned by the speaker US (e.g., a PC, laptop, smartphone, or tablet device). The speaker information here may be, for example, identification information that can identify the speaker US, the name of the speaker US, speaker identification (ID), etc.

[0032] The similarity calculation model selection unit 115, which is an example of a selection unit, calculates (evaluates) the similarity between the noise type associated with each feature of a plurality of registered speakers registered in the registered speaker database DB2 and the noise type of noise extracted from the registered speech data of the speaker US. Based on the calculated similarity, the similarity calculation model selection unit 115 selects a similarity calculation model (an example of a similarity calculation model) stored in the similarity calculation model database DB3.

[0033] The similarity calculation model selected here is a model that is more suitable or optimal for calculating the similarity between the feature amount of the speaker US and the feature amount of any of the registered speakers.

[0034] The similarity calculation model selection unit 115 refers to correspondence lists LST1, LST2 (see Figures 5 and 6) that associate information on noise types corresponding to the registered voice data of speaker US, information on noise types corresponding to multiple registered speakers, and information on similarity calculation models selected based on the similarities of these noise types, selects one selection model from each of multiple selection models (similarity models), and outputs it to each of the reliability calculation unit 116 and the authentication unit 117.

[0035] Confidence calculation unit 116, as an example of a confidence calculation unit, calculates (evaluates) a confidence (score) indicating the likelihood of the identification result of speaker US based on the similarity calculated by authentication unit 117. Confidence calculation unit 116 outputs information on the calculated confidence to authentication unit 117 based on the similarity calculated by authentication unit 117, information on the noise type used in the similarity calculation process, a similarity calculation model, etc. Note that the confidence calculation process by confidence calculation unit 116 is not essential and may be omitted.

[0036] The authentication unit 117, which is an example of a calculation unit, acquires the features of the speaker US output from the feature extraction unit 112, and acquires the features of each of the multiple registered speakers registered in the registered speaker database DB2. The authentication unit 117 also acquires the selected models of the correspondence lists LST1 and LST2 output from the similarity calculation model selection unit 115.

[0037] The authentication unit 117 calculates the similarity between the feature amounts of each of the multiple registered speakers and the feature amounts of the speaker US using a similarity calculation model based on the correspondence lists LST1 and LST2. The authentication unit 117 identifies the speaker US based on the calculated similarity. The authentication unit 117 generates an authentication result screen SC based on the speaker information of the identified speaker US and transmits it to the monitor MN.

[0038] The memory 12 includes, for example, a random access memory (hereinafter referred to as "RAM") as a work memory used when executing each process of the processor 11, and a read only memory (hereinafter referred to as "ROM") that stores programs and data that define the operation of the processor 11. The RAM temporarily stores data or information generated or acquired by the processor 11. The ROM stores programs that define the operation of the processor 11.

[0039] The feature extraction model database DB1 is a so-called storage device, and is configured using a storage medium such as a flash memory, a hard disk drive (hereinafter referred to as "HDD"), or a solid state drive (hereinafter referred to as "SSD"). The feature extraction model database DB1 stores a feature extraction model that can detect speech periods of a speaker US from enrollment voice data or authentication voice data and extract features of this speaker US. The feature extraction model is, for example, a learning model generated by learning using deep learning or the like.

[0040] The registered speaker database DB2 is a so-called storage, and is configured using a storage medium such as a flash memory, HDD, SSD, etc. The registered speaker database DB2 stores the feature amounts of each of a plurality of registered speakers registered in advance, information on the noise type of the noise included in the registered speech data from which the feature amounts were extracted, and the registered speaker information, in association with each other.

[0041] The similarity calculation model database DB3 is a so-called storage, and is configured using a storage medium such as a flash memory, HDD, or SSD. The similarity calculation model database DB3 stores a similarity calculation model capable of calculating the similarity between two feature quantities. The similarity calculation model is, for example, a learning model generated by learning using deep learning or the like.

[0042] For example, the similarity calculation model is a model that pre-learns and stores dimensions that are likely to reveal individual characteristics in order to calculate the similarity between two multidimensional vectors with high accuracy. Note that the method of calculating similarity using the model is one example of a method for calculating similarity between vectors, and previously mentioned techniques such as Euclidean distance and cosine similarity may also be used.

[0043] The noise-similarity calculation model correspondence list DB4 is a so-called storage, and is configured using a storage medium such as a flash memory, HDD, SSD, etc. The noise-similarity calculation model correspondence list DB4 stores the similarity calculation model used in the similarity calculation process for each combination of noise types.

[0044] The monitor MN is configured using a display such as a Liquid Crystal Display (LCD) or an organic electroluminescence (EL) display, etc. The monitor MN displays the authentication result screen SC output from the terminal device P1.

[0045] The authentication result screen SC is a screen that notifies an administrator (e.g., a person watching the monitor MN) of the speaker authentication result, and includes authentication result information such as "The voice matches that of XX XX" and reliability information such as "High reliability." The authentication result screen SC may also include other registered speaker information (e.g., a facial image, etc.). Furthermore, the authentication result screen SC does not necessarily have to include reliability information.

[0046] The network NW connects the terminal device P1 and the noise determination device P2 so that data communication can be performed between them. Note that the noise determination device P2 may not only be connected to the terminal device P1 via the network NW, but may also be a part of the terminal device P1.

[0047] The noise determination device P2 acquires the enrollment voice data of the speaker US transmitted from the terminal device P1 or the noise extracted from the authentication voice data. The noise determination device P2 determines the type of noise based on the acquired noise. The noise determination device P2 transmits information on the noise type to the terminal device P1.

[0048] Next, the operation procedure of the terminal device P1 will be described with reference to Fig. 3. Fig. 3 is a flowchart showing an example of the operation procedure of the terminal device P1 in the embodiment.

[0049] The terminal device P1 acquires voice data from the microphone MK (St11). Note that the microphone MK may be a microphone provided in, for example, a PC, a laptop PC, a smartphone, or a tablet terminal.

[0050] The terminal device P1 determines whether the control command associated with the voice data is a control command requesting registration in the registered speaker database DB2 (St12).

[0051] In the process of step St12, if the control command is a control command requesting registration in the registration speaker database DB2, the terminal device P1 determines to newly register the features of the speaker US in the registration speaker database DB2 (St12, YES), and extracts noise included in the non-speech section of the voice data (enrollment voice data) (St13). The noise here refers to noise included in the voice data (enrollment voice data or authentication voice data), such as ambient sound or noise around the time the voice of the speaker US is picked up.

[0052] On the other hand, in the processing of step St12, if the control command is not a control command requesting registration in the registered speaker database DB2 but a control command requesting speaker authentication, the terminal device P1 determines not to newly register the features of the speaker US in the registered speaker database DB2 (St12, NO), and extracts noise contained in the non-speech section of the voice data (authentication voice data) (St14).

[0053] The terminal device P1 associates the extracted noise with a control command requesting a determination of the noise type, and transmits the extracted noise to the noise determination device P2. The terminal device P1 executes a noise type determination process by acquiring the information on the noise type (i.e., the determination result) transmitted from the noise determination device P2 (St15).

[0054] The terminal device P1 extracts features indicative of the individuality of the speaker US from the speech sections of the enrollment voice data (St16). Here, the features extracted from the speech sections of the enrollment voice data and the authentication voice data include features indicative of the individuality of the speaker US and features of noise.

[0055] The terminal device P1 associates the feature amount of the speaker US extracted from the enrollment voice data with the noise type information and the speaker information, and registers them in the enrollment speaker database DB2 (St17).

[0056] The terminal device P1 further extracts noise extracted from the non-utterance section of the authentication voice data. The terminal device P1 associates the extracted noise with a control command requesting a noise type determination, and transmits the control command to the noise determination device P2. The terminal device P1 acquires the noise type information (i.e., the determination result) transmitted from the noise determination device P2, and executes a noise type determination process (St18).

[0057] The terminal device P1 extracts features of the speaker US from the speech section of the authentication voice data of the speaker US (St19). The terminal device P1 also acquires features of each of the multiple registered speakers registered in the registered speaker database DB2 (St20) and executes speaker authentication processing (St21).

[0058] Next, the speaker authentication procedure shown in step St21 in Fig. 3 will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the speaker authentication procedure of the terminal device P1 in the embodiment.

[0059] The terminal device P1 selects, for each registered speaker, a similarity calculation model suitable for calculating the similarity between the feature of the speaker US and each of the multiple registered speakers, based on the noise type of the speech data of the speaker US and the noise type of each of the multiple registered speakers registered in the registered speaker database DB2. The terminal device P1 refers to correspondence lists LST1 and LST2 (see FIGS. 5 and 6) that associate the noise type of the speech data of the speaker US, the noise type of each of the multiple registered speakers, and the selected similarity calculation model.

[0060] Based on the referenced correspondence lists LST1 and LST2, the terminal device P1 reads from the similarity calculation model database DB3 a similarity calculation model used to determine the similarity between the features of the speaker US and the features of any one of the multiple registered speakers (St211).

[0061] The terminal device P1 uses the similarity calculation model to calculate the similarity between the feature of the speech data of the speaker US and the feature of one of the registered speakers registered in the registered speaker database DB2 (St212). The terminal device P1 repeatedly executes the process of step St212 until it has calculated the similarity between the feature of the speech data of the speaker US and the feature of all registered speakers registered in the registered speaker database DB2.

[0062] The terminal device P1 determines whether any of the calculated similarities is equal to or greater than a threshold value (St213).

[0063] If the terminal device P1 determines in the processing of step St213 that there is a similarity equal to or greater than a threshold among the calculated similarities (St213, YES), it identifies the speaker US based on the registered speaker information corresponding to the similarity determined to be equal to or greater than the threshold (St214). Note that, if there are multiple similarities determined to be equal to or greater than the threshold, the terminal device P1 may identify the speaker US based on the registered speaker information corresponding to the highest calculated similarity.

[0064] If it is determined in the process of step St213 that none of the calculated similarities is equal to or greater than the threshold (St213, NO), the terminal device P1 determines that the speaker US cannot be identified (St215).

[0065] The terminal device P1 generates an authentication result screen SC based on the registered speaker information of the identified speaker US. The terminal device P1 outputs the generated authentication result screen SC to the monitor MN for display (St216).

[0066] As described above, the terminal device P1 registers speaker information, features of the speaker US, and information on the noise types included in the speech sections of the speaker US in association with each other during voice registration. This allows the terminal device P1 to select a similarity calculation model appropriate for each noise type, even if the noise types included in the features during voice registration differ from the noise types included in the features during voice authentication. Therefore, by using the selected similarity calculation model, the terminal device P1 can more accurately determine the similarity between the features of the speaker US, which contain different noises, and the features of the registered speaker, thereby more effectively suppressing a decrease in speaker authentication accuracy due to noise included in the authentication voice data.

[0067] Furthermore, the terminal device P1 calculates and displays a reliability indicating the likelihood of the speaker identified by the speaker authentication process, as indicated by the calculated similarity, based on the similarity calculation model used to calculate the similarity. This allows the terminal device P1 to present the reliability of the speaker authentication result to the administrator watching the monitor MN. Therefore, by presenting the reliability, the terminal device P1 can inform the administrator that there is no similarity calculation model suitable for calculating the similarity, and that speaker authentication was performed using a similarity calculation model "general-purpose model" (described later).

[0068] Examples of correspondence lists LST1 and LST2 will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a diagram illustrating an example of correspondence list LST1 when the noise type at the time of voice registration is the same as the noise type at the time of voice authentication. Fig. 6 is a diagram illustrating an example of correspondence list LST2 when the noise type at the time of voice registration is different from the noise type at the time of voice authentication.

[0069] In addition, in Figures 5 and 6, for ease of explanation, an example is described in which correspondence lists LST1 and LST2 are referenced based on noise type information associated with the features of any one of the registered speakers registered in the registered speaker database DB2 and the noise type of noise contained in the authentication voice data of speaker US, who is the target of speaker authentication.

[0070] In the reference example of the correspondence list LST1 shown in FIG. 5, the voice data at the time of voice registration and voice authentication each contain noise that falls under the same noise type, "in-store noise."

[0071] The terminal device P1 transmits noise extracted from the authentication voice data of the speaker US transmitted from the microphone MK to the noise determination device P2. The terminal device P1 selects one selection model from among a plurality of selection models (similarity models) by referring to a predefined correspondence list LST1 based on the noise type determination result transmitted from the noise determination device P2 and the noise type information of the registered speaker registered in the registered speaker database DB2.

[0072] The correspondence list LST1 is data that associates the noise type determination result "noise determination result 1" of the noise contained in the authentication voice data of speaker US, the noise type determination result "noise determination result 2" of the registered speaker registered in the registered speaker database DB2, and the similarity calculation model "selected model" selected based on these two noise types.

[0073] The noise type determination result "noise determination result 1" indicates, for example, the noise determination result for the voice during authentication.

[0074] The noise type determination result "noise determination result 2" indicates, for example, the noise determination result of the registered voice.

[0075] It is sufficient that one or more pieces of noise type information are registered in the registered speaker database DB2. In other words, if there are multiple noise type candidates, multiple pieces of information may be stored. Furthermore, information on the determination probability, which indicates the reliability of each of the determination results for the multiple noise types, is not essential and may be omitted.

[0076] The similarity calculation model "selected model" includes, for example, the similarity calculation models "Model A," "Model B," "Model C," and "Model Z" selected in correspondence with the noise type determination result "Noise Determination Result 1" and the noise type determination result "Noise Determination Result 2."

[0077] For example, the similarity calculation model "Model A" is a similarity calculation model that is determined to be optimal for calculating the similarity between the features of speaker US and the features of a registered speaker when the noise type information of the noise included in the features of speaker US is "Noise A" and the noise type information of the noise included in the features of a registered speaker is "Noise A."

[0078] In addition, if the similarity calculation model selection unit 115 determines that there is no similarity calculation model suitable for the similarity calculation process based on a combination of information on the noise type of noise contained in the features of speaker US and information on the noise type of noise contained in the features of the registered speakers, it selects "Model Z," which is a general-purpose similarity calculation model.

[0079] The similarity calculation model selection unit 115 selects a similarity calculation model database from the similarity calculation model database DB3 based on the referenced correspondence list LST1. For example, in the example shown in Fig. 5, the similarity calculation model selection unit 115 selects the similarity calculation model "Model A".

[0080] Next, in the reference example of correspondence list LST2 shown in Figure 6, the authentication voice data of speaker US at the time of voice registration includes noise that corresponds to the noise type "in-store noise." Also, the enrollment voice data of the enrolled speaker at the time of voice authentication includes noise that corresponds to the noise type "outdoor noise," which is different from the authentication voice data of speaker US.

[0081] The terminal device P1 transmits noise extracted from the authentication voice data of the speaker US transmitted from the microphone MK to the noise determination device P2. The terminal device P1 selects one selection model from each of a plurality of selection models (similarity models) by referring to a predefined correspondence list LST2 based on the noise type determination result transmitted from the noise determination device P2 and the noise type information of the registered speaker registered in the registered speaker database DB2.

[0082] The correspondence list LST2 is data that associates the noise type determination result "noise determination result 3" of the noise contained in the authentication voice data of speaker US, the noise type determination result "noise determination result 4" of the registered speaker registered in the registered speaker database DB2, and the similarity calculation model "selected model" selected based on these two noise types.

[0083] The noise type determination result "noise determination result 3" indicates the determination result of the noise type extracted from the voice data of the registered speaker.

[0084] The noise type determination result "noise determination result 4" indicates the determination result of the noise type information of the registered speaker registered in the registered speaker database DB2.

[0085] It is sufficient that one or more pieces of noise type information are registered in the registered speaker database DB2. In other words, if there are multiple noise type candidates, multiple pieces of information may be stored. Furthermore, information on the determination probability, which indicates the reliability corresponding to the determination result of each noise type, is not essential and may be omitted.

[0086] The similarity calculation model "selected model" includes the similarity calculation models "Model G," "Model H," "Model I," and "Model Z" selected in response to the noise type determination result "Noise Determination Result 3" and the noise type determination result "Noise Determination Result 4."

[0087] For example, the similarity calculation model "Model G" is a similarity calculation model that is determined to be optimal for calculating the similarity between the features of speaker US and the features of a registered speaker when the noise type information of the noise included in the authentication voice data is "Noise A" and the noise type information of the registered speaker is "Noise D."

[0088] In addition, if the similarity calculation model selection unit 115 determines that there is no similarity calculation model suitable for the similarity calculation process based on a combination of information on the noise type of noise contained in the feature of speaker US and information on the noise type of the registered speaker, it selects the general-purpose similarity calculation model "Model Z."

[0089] The similarity calculation model selection unit 115 selects a similarity calculation model database from the similarity calculation model database DB3 based on the referenced correspondence list LST2. For example, in the example shown in Fig. 6, the similarity calculation model selection unit 115 selects the similarity calculation model "Model E".

[0090] As described above, the terminal device P1 can select an optimal similarity calculation model for the similarity calculation process between two features (features of speaker US and features of registered speakers) based on the combination of noise types contained in the features of speaker US, which are the targets of similarity calculation, and the features of registered speakers. This allows the terminal device P1 to select an optimal similarity calculation model for calculating the similarity between two features, even if the noise types contained in the features at the time of voice registration and the features at the time of voice authentication change. In other words, the terminal device P1 can more effectively suppress a decrease in speaker authentication accuracy due to noise contained in voice data. When multiple candidate conditions exist during noise determination, the terminal device P1 may calculate similarities using similarity calculation models corresponding to each candidate, and then calculate an average value to use as the similarity.

[0091] The reliability of the similarity calculation model used in the similarity calculation process will be described with reference to Fig. 7. Fig. 7 is a diagram for explaining an example of calculating the reliability.

[0092] Although FIG. 7 shows an example in which the reliability is calculated using two indices, "high" and "low," the reliability may be calculated using a numerical value, for example, from 0 (zero) to 100.

[0093] In addition, for ease of understanding, Figure 7 will explain an example of calculating the reliability when the registration voice data during voice registration and the authentication voice data during voice authentication are of the same noise type, as explained in Figure 5.

[0094] The reliability calculation unit 116 determines the reliability of the similarity calculated by the authentication unit 117. Here, the reliability calculation unit 116 determines the reliability based on whether the similarity calculation model used to calculate the similarity is a similarity calculation model based on a known noise type, the determination probability of the noise type, etc.

[0095] In the example shown in "Case 1," the noise type determination results based on noise are that the noise type is determined to be "outdoor noise" with a probability of "90%, the noise type is determined to be "in-store music" with a probability of "6%, and the noise type is determined to be "unknown noise" with a probability of "4%." The noise type determination results shown in "Case 1" also indicate that the noise types "outdoor noise" and "in-store noise" are known noises, and the noise type "unknown noise" is unknown noise.

[0096] The similarity calculation model selection unit 115 selects the similarity calculation model "outdoor wind noise model" based on the noise type determination result of the noise. The authentication unit 117 uses the similarity calculation model "outdoor noise model" to calculate the similarity between the feature of the speaker US and the feature of any registered speaker registered in the registered speaker database DB2.

[0097] The reliability calculation unit 116 determines the reliability of the similarity calculated by the authentication unit 117. Here, the reliability calculation unit 116 calculates the reliability as "high" because the similarity calculation model "outdoor noise model" is a similarity calculation model based on the known noise type "outdoor wind noise model" and the noise determination probability is "90%."

[0098] In addition, if the reliability calculation unit 116 determines that the noise determination probability is equal to or greater than a predetermined probability (e.g., 85%, 90%, etc.), it calculates the reliability to be "high," and if it determines that the noise determination probability is not equal to or greater than the predetermined probability, it calculates the reliability to be "low."

[0099] In the example shown in "Case 2," the noise type determination results based on noise are that the noise type is determined to be "outdoor wind noise" with a probability of "48%, the noise type is determined to be "unknown noise" with a probability of "39%, and the noise type is determined to be "in-store music" with a probability of "13%." The noise type determination results shown in "Case 2" also indicate that the noise types "outdoor wind noise" and "in-store music" are known noises, and the noise type "unknown noise" is an unknown noise.

[0100] The similarity calculation model selection unit 115 selects the similarity calculation model "outdoor wind noise model" based on the noise type determination result of the noise. The authentication unit 117 uses the similarity calculation model "outdoor noise model" to calculate the similarity between the feature of the speaker US and the feature of any registered speaker registered in the registered speaker database DB2.

[0101] The reliability calculation unit 116 determines the reliability of the similarity calculated by the authentication unit 117. Here, the reliability calculation unit 116 calculates the reliability as "low" because the similarity calculation model "outdoor noise model" is a similarity calculation model based on the known noise type "outdoor wind noise model" and the noise determination probability is "48%."

[0102] In the example shown in "Case 3," the noise type determination results based on noise are that the noise type is determined to be "unknown noise" with a probability of "55%, "outdoor wind noise" with a probability of "28%, and "in-store music" with a probability of "17%." The noise type determination results shown in "Case 3" indicate that the noise types "outdoor wind noise" and "in-store noise" are known noises, while the noise type "unknown noise" is an unknown noise.

[0103] The similarity calculation model selection unit 115 selects the similarity calculation model "general-purpose model" based on the noise type determination result of the noise. The authentication unit 117 uses the similarity calculation model "general-purpose model" to calculate the similarity between the feature of the speaker US and the feature of any registered speaker registered in the registered speaker database DB2.

[0104] The reliability calculation unit 116 determines the reliability of the similarity calculated by the authentication unit 117. Here, the reliability calculation unit 116 calculates the reliability as "low" because the similarity calculation model "general-purpose model" used to calculate the similarity is unknown noise and the noise determination probability is "55%."

[0105] The reliability calculation unit 116 determines the reliability of the similarity calculated by the authentication unit 117. Here, the reliability calculation unit 116 calculates the reliability as "low" because the similarity calculation model "outdoor noise model" is a similarity calculation model based on the known noise type "outdoor wind noise model" and the noise determination probability is "48%."

[0106] As described above, the terminal device P1 according to the embodiment includes a communication unit 10 (an example of an acquisition unit) that acquires voice data, a noise extraction unit 111 and a feature extraction unit 112 (an example of a detection unit) that detect speech intervals in which the speaker is speaking and non-speech intervals in which the speaker is not speaking from the voice data, a noise extraction unit 111 and a feature extraction unit 112 (an example of an extraction unit) that extract features of the speech intervals (an example of speech features) and noise included in the non-speech intervals of the voice data, a similarity calculation model selection unit 115 (an example of a selection unit) that selects one of a plurality of similarity calculation models (an example of similarity calculation models) based on the extracted noise and noise associated with features of a plurality of registered speakers that have been registered in advance (an example of registered features), and an authentication unit 117 that uses the selected similarity calculation model to compare the features of the speaker US with the features of the registered speakers to authenticate the speaker.

[0107] As a result, the terminal device P1 according to the embodiment can select a similarity calculation model suitable for speaker authentication based on a combination of noise types contained in the features of the speaker US and the features of the registered speakers. That is, even if the noise types contained in the features at the time of voice registration differ from the noise types contained in the features at the time of voice authentication, the terminal device P1 can select a similarity calculation model more suitable for speaker authentication. Therefore, the terminal device P1 can more effectively suppress a decrease in speaker authentication accuracy caused by noise contained in voice data.

[0108] Furthermore, the communication unit 10 in the terminal device P1 according to the embodiment further acquires noise type information of the extracted noise. The similarity calculation model selection unit 115 selects a similarity calculation model based on the acquired noise type information of the speaker and the noise type information of the registered speakers. This allows the terminal device P1 according to the embodiment to select a similarity calculation model that is more suitable for the similarity calculation process of two feature amounts (the feature amounts of speaker US and the feature amounts of the registered speakers) based on the combination of noise types included in the feature amounts of speaker US and the feature amounts of the registered speakers.

[0109] Furthermore, the terminal device P1 according to the embodiment further includes an authentication unit 117 (an example of a calculation unit) that calculates similarities between the features of the speaker US and the features of registered speakers. The authentication unit 117 authenticates the speaker US based on the calculated similarities. This allows the terminal device P1 according to the embodiment to perform speaker authentication using similarities between the features of the speaker US and the features of multiple registered speakers who have been registered in advance.

[0110] Moreover, the terminal device P1 according to the embodiment further includes a reliability calculation unit 116 (an example of a reliability calculation unit) that calculates the reliability of the similarity. The communication unit 10 further acquires noise type information of the extracted noise and a score for which the noise is the noise type. The reliability calculation unit 116 calculates the reliability of the similarity based on the score. As a result, the terminal device P1 according to the embodiment can calculate the reliability of the speaker authentication result by calculating the reliability corresponding to the similarity.

[0111] Furthermore, the authentication unit 117 in the terminal device P1 according to the embodiment identifies a registered speaker whose similarity is equal to or greater than a threshold as the speaker US. This allows the terminal device P1 according to the embodiment to perform speaker authentication using the similarity between the feature amounts of a plurality of registered speakers who have been registered in advance and the feature amounts of the speaker US.

[0112] Furthermore, the authentication unit 117 in the terminal device P1 according to the embodiment generates and outputs an authentication result screen SC including information about registered speakers whose similarity is equal to or greater than a threshold value, thereby allowing the terminal device P1 according to the embodiment to present the speaker authentication result to the speaker US or an administrator.

[0113] Furthermore, when it is determined that the calculated similarities are not equal to or greater than a threshold, the authentication unit 117 in the terminal device P1 according to the embodiment determines that the speaker US cannot be identified. This allows the terminal device P1 according to the embodiment to more effectively prevent a decrease in speaker authentication accuracy and more effectively prevent erroneous authentication of the speaker US.

[0114] Furthermore, the authentication unit 117 in the terminal device P1 according to the embodiment generates and outputs an authentication result screen SC including information on registered speakers whose similarity is equal to or greater than a threshold and information on the calculated reliability. As a result, the terminal device P1 according to the embodiment can prompt the administrator to confirm whether the speaker authentication result is reliable by displaying the speaker authentication result and the reliability of the speaker authentication result.

[0115] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention.

[0116] This application is based on a Japanese patent application (Patent Application No. 2022-045390) filed on March 22, 2022, the contents of which are incorporated herein by reference. [Industrial Applicability]

[0117] The present disclosure is useful for a voice authentication device and a voice authentication method that suppress a decrease in speaker authentication accuracy caused by changes in environmental noise. [Explanation of symbols]

[0118] 10. Communications Department 11 processors 12 Memory 100 Voice Authentication System 111 Noise Extraction Unit 112 Feature Extraction Unit 113 Noise determination unit 114 Speaker Registration Department 115 Similarity calculation model selection unit 116 Reliability calculation unit 117 Authentication Department DB1 Feature Extraction Model Database DB2 registered speaker database DB3 Similarity Calculation Model Database DB4 noise-similarity calculation model compatibility list MK Microphone MN Monitor NW Network P1 terminal equipment P2 noise detector SC authentication result screen US speakers

Claims

1. an acquisition unit that acquires voice data; a detection unit that detects, from the audio data, speech sections in which a speaker is speaking and non-speech sections in which the speaker is not speaking; an extracting unit that extracts speech features of the speech section and noise included in the non-speech section of the audio data; a selection unit that selects one of a plurality of similarity calculation models based on the extracted noise and noise associated with registered features of a plurality of registered speakers that have been registered in advance; and an authentication unit that uses the selected similarity calculation model to compare the speech features of the speaker with the registered features of the registered speakers to authenticate the speaker. Voice authentication device.

2. The acquisition unit further acquires noise type information of the extracted noise, the selection unit selects the similarity calculation model based on the acquired noise type information of the speaker and the noise type information of the registered speaker. The voice authentication device according to claim 1 .

3. a calculation unit that calculates a similarity between the speech feature of the speaker and the registered feature of the registered speaker, the authentication unit authenticates the speaker based on the calculated similarities. The voice authentication device according to claim 1 .

4. a reliability calculation unit that calculates the reliability of the similarity; the acquiring unit further acquires information on the noise type of the extracted noise and a score for the noise as the noise type; the reliability calculation unit calculates the reliability of the similarity based on the score. The voice authentication device according to claim 3.

5. The authentication unit identifies a registered speaker whose similarity is equal to or greater than a threshold as the speaker. The voice authentication device according to claim 3.

6. the authentication unit generates and outputs an authentication result screen including information about the registered speakers whose similarity is equal to or greater than the threshold. The voice authentication device according to claim 5.

7. When the authentication unit determines that the calculated similarities are not equal to or greater than a threshold, the authentication unit determines that the speaker cannot be identified. The voice authentication device according to claim 3.

8. the authentication unit identifies a registered speaker whose similarity is equal to or greater than a threshold as the speaker, and generates and outputs an authentication result screen including information on the registered speaker whose similarity is equal to or greater than the threshold and information on the calculated reliability. The voice authentication device according to claim 4.

9. A voice authentication method performed by a terminal device, Acquires audio data, detecting a speech section in which a speaker is speaking and a non-speech section in which the speaker is not speaking from the audio data; extracting a speech feature of the speech section and noise included in the non-speech section of the audio data; selecting one of a plurality of similarity calculation models based on the extracted noise and noise associated with the registered features of a plurality of registered speakers that have been registered in advance; using the selected similarity calculation model, the speech feature of the speaker is compared with the registered feature of the registered speaker to authenticate the speaker; Voice authentication method.

Citation Information

Patent Citations

  • Voice recognition equipment

    JP1987042198A

  • Speech recognizing method

    JP1993073090A

  • Pattern matching system

    JP1995036477A

  • On-board voice recognition system

    JP2006003400A

  • Voice recognition device, voice recognition system and voice recognition method

    JP2008250059A