Authentication system and authentication method
The authentication system improves user convenience by adjusting authentication conditions based on speech duration and quality, addressing the inconvenience of fixed speech times in existing voiceprint systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2026-04-03
AI Technical Summary
Existing voiceprint authentication systems require users to speak for a predetermined time during registration and authentication, which is inconvenient for users.
An authentication system that determines authentication conditions based on the length of the user's speech during registration, allowing for shorter or longer speech durations depending on the quality of the speech, thereby improving user convenience.
The system enhances user convenience by allowing flexible speech durations during authentication while maintaining high authentication accuracy.
Smart Images

Figure 0007839991000001 
Figure 0007839991000002 
Figure 0007839991000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to an authentication system and an authentication method.
Background Art
[0002] Patent Document 1 discloses a communication device that registers voiceprint data for voiceprint authentication from the received voice during a call. The communication device acquires the received voice, acquires the telephone number of the calling party, and extracts voiceprint data from the acquired received voice. Next, the communication device measures the acquisition time of the received voice. The communication device determines whether the total acquisition time length of at least one or more voiceprint data corresponding to the same telephone number as the acquired telephone number in the telephone directory is longer than the time required for voiceprint verification. When the communication device determines that the total acquisition time of the voiceprint data is longer than the time required for voiceprint verification, the communication device associates the acquired telephone number with the voiceprint data and stores them in the storage unit.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In Patent Document 1, when the total acquisition time length of the speaker's voiceprint data becomes a predetermined value or more, the voiceprint data is linked to the speaker's telephone number and registered in the database. That is, the communication device disclosed in Patent Document 1 always requires voiceprint data for registration used in voiceprint authentication for a total acquisition time length of a predetermined value or more. For this reason, the user needs to speak for a time of a predetermined value or more for registration of voiceprint data, and thus also needs to speak for a similar time during voiceprint authentication. Therefore, an improvement for better user convenience is expected.
[0005] This disclosure was devised in light of the aforementioned conventional circumstances and aims to improve user convenience by determining the speech duration at the time of authentication based on the total length of the user's speech acquired at the time of registration. [Means for solving the problem]
[0006] This disclosure includes: an acquisition unit that acquires an audio signal of a speaker's utterance; a detection unit that detects a first utterance section spoken by the speaker from the acquired audio signal and a second utterance section spoken by the speaker from the audio signals in a database in which the audio signals of multiple speakers are registered; a determination unit that compares the first audio signal of the first utterance section and the second audio signal of the second utterance section and determines authentication conditions for authentication using the first audio signal based on the length of the second audio signal of the second utterance section or the number of sounds included in the second utterance section; and an authentication unit that performs authentication of the speaker based on the determined authentication conditions. The authentication unit starts authentication when the length of the second utterance is greater than or equal to a first predetermined value and the length of the first utterance is greater than or equal to a second predetermined value, and starts authentication when the length of the second utterance is less than the first predetermined value and the length of the first utterance is greater than or equal to a third predetermined value which is greater than the second predetermined value. We provide an authentication system.
[0007] Furthermore, this disclosure relates to an authentication method performed by one or more computers, which involves: acquiring an audio signal of a speaker's utterance; detecting a first utterance section spoken by the speaker from the acquired audio signal; detecting a second utterance section spoken by the speaker from the audio signals in a database in which the audio signals of multiple speakers are registered; comparing the first audio signal of the first utterance section with the second audio signal of the second utterance section; determining authentication conditions for authentication using the first audio signal based on the length of the second audio signal of the second utterance section or the number of sounds contained in the second utterance section; and performing authentication of the speaker based on the determined authentication conditions. stomach , Authentication is initiated when the length of the second utterance is greater than or equal to a first predetermined value, and the length of the first utterance is greater than or equal to a second predetermined value; and authentication is initiated when the length of the second utterance is less than the first predetermined value, and the length of the first utterance is greater than or equal to a third predetermined value that is greater than the second predetermined value. Provide an authentication method.
[0008] These comprehensive or specific embodiments may be implemented as systems, devices, methods, integrated circuits, computer programs, or recording media, or as any combination of systems, devices, methods, integrated circuits, computer programs, and recording media. [Effects of the Invention]
[0009] According to this disclosure, the speech duration at the time of authentication can be determined based on the total length of the user's speech acquired during registration, thereby improving user convenience. [Brief explanation of the drawing]
[0010] [Figure 1] A diagram showing an example of a use case for the authentication system according to this embodiment. [Figure 2] Block diagram showing an example of the internal configuration of the authentication analysis device according to this embodiment. [Figure 3] Flowchart for the registration process of speech audio signals for registration. [Figure 4] A diagram showing an example of setting authentication conditions based on utterance length. [Figure 5] A diagram showing an example of setting authentication conditions based on utterance length and number of syllables. [Figure 6] This diagram illustrates an example of setting speech content as authentication conditions based on the quality of the recorded voice signal. [Figure 7] This diagram illustrates an example of an operator performing identity verification based on authentication text displayed on the screen. [Figure 8] This diagram illustrates an example of performing identity verification based on the identity verification text displayed on the user's phone terminal. [Figure 9] This diagram illustrates an example of resetting the required time for authentication conditions based on the measurement results of the sound collection conditions during authentication, after the authentication conditions have been set. [Figure 10] This diagram illustrates an example of resetting the authentication threshold based on the measurement results of the sound collection conditions during authentication, after the authentication conditions have been set. [Figure 11] Flowchart for speaker authentication process [Figure 12] This diagram illustrates an example of restricting actions after successful authentication based on the quality of the registered audio signal. [Figure 13] Flowchart for processing operational restrictions based on the quality of registered audio signals [Modes for carrying out the invention]
[0011] Hereinafter, embodiments specifically disclosing the authentication system and authentication method according to the present disclosure will be described in detail with appropriate reference to the drawings. However, detailed descriptions that are more than necessary may be omitted. For example, detailed descriptions of well-known matters and duplicate descriptions of substantially the same configurations may be omitted. This is to avoid making the following description unnecessarily redundant and to facilitate the understanding of those skilled in the art. Note that the attached drawings and the following description are provided for those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.
[0012] First, referring to FIG. 1, the use case of the authentication system according to this embodiment will be described. FIG. 1 is a diagram showing an example of the use case of the authentication system according to this embodiment. The authentication system 100 acquires the voice signal or voice data of a person to be authenticated by voice (in the example shown in FIG. 1, user US), and collates the acquired voice signal or voice data with the voice signal or voice data of the speaker registered (stored) in advance in a storage (in the example shown in FIG. 1, the registered speaker database DB). The authentication system 100 evaluates the similarity between the voice signal or voice data collected from the user US to be authenticated and the voice signal or voice data registered in the storage based on the collation result, and authenticates the user US based on the evaluated similarity.
[0013] The authentication system 100 according to Embodiment 1 includes an operator-side communication terminal OP1 as an example of a sound collection device, an authentication analysis device P1, a registered speaker database DB, and a display DP as an example of an output device. Note that the authentication analysis device P1 and the display DP may be integrally configured. Note that the operator-side communication terminal OP1 may be replaced with an automatic voice device, and in this case, the automatic voice device may be integrally configured with the authentication analysis device P1.
[0014] Note that the authentication system 100 shown in FIG. 1 shows an example of being used for authenticating a speaker (user US) in a call center, and authenticates the user US using voice data obtained by collecting the voice of the user US speaking with the operator OP. The authentication system 100 shown in FIG. 1 further includes a user-side call terminal UP1 and a network NW. Needless to say, the overall configuration of the authentication system 100 is not limited to the example shown in FIG. 1.
[0015] The user-side call terminal UP1 is wirelessly communicably connected to the operator-side call terminal OP1 via the network NW. Here, the wireless communication is, for example, network communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark).
[0016] The user-side call terminal UP1 is constituted by, for example, a notebook PC, a tablet terminal, a smartphone, or a telephone. The user-side call terminal UP1 is a sound collection device provided with a microphone (not shown), collects the voice of the user US and converts it into a voice signal, and transmits this converted voice signal to the operator-side call terminal OP1 via the network NW. Further, the user-side call terminal UP1 acquires the voice signal of the voice of the operator OP transmitted from the operator-side call terminal OP1 and outputs it from a speaker (not shown).
[0017] The network NW is, for example, an IP (Internet Protocol) network or a telephone network, and enables the transmission and reception of voice signals between the user-side call terminal UP1 and the operator-side call terminal OP1. The transmission and reception of data are performed by wired communication or wireless communication.
[0018] The operator-side call terminal OP1 is respectively connected to the user-side call terminal UP1 and the authentication analysis device P1 so as to be able to transmit and receive data by wired communication or wireless communication, and performs the transmission and reception of voice signals.
[0019] The operator-side call terminal OP1 is composed of, for example, a notebook PC, tablet, smartphone, or telephone. The operator-side call terminal OP1 acquires an audio signal based on the user US's speech transmitted from the user-side call terminal UP1 via the network NW and transmits it to the authentication analysis device P1. If the operator-side call terminal OP1 acquires an audio signal that includes both the user US's speech and the operator OP's speech, it may separate the audio signal based on the user US's speech from the audio signal based on the operator OP's speech based on audio parameters such as the sound pressure level and frequency band of the audio signal of the operator-side call terminal OP1. After separation, the operator-side call terminal OP1 extracts only the audio signal based on the user US's speech and transmits it to the authentication analysis device P1.
[0020] Furthermore, the operator-side call terminal OP1 may be connected to each of the multiple user-side call terminals in a communication-enabled manner, and may simultaneously acquire voice signals from each of the multiple user-side call terminals. The operator-side call terminal OP1 transmits the acquired voice signals to the authentication analysis device P1. This allows the authentication system 100 to simultaneously perform voice authentication processing and voice analysis processing for each of the multiple users.
[0021] Furthermore, the operator-side call terminal OP1 may simultaneously acquire audio signals containing the individual speech voices of multiple users. The operator-side call terminal OP1 extracts a user-specific audio signal from each of the multiple user audio signals acquired via the network NW and transmits each user-specific audio signal to the authentication analysis device P1. In such cases, the operator-side call terminal OP1 may analyze the audio signals of multiple users and separate and extract the audio signals for each user based on audio parameters such as sound pressure level and frequency band. If the audio signals are collected by an array microphone or the like, the operator-side call terminal OP1 may separate and extract the audio signals for each user based on the direction of arrival of the utterance. This allows the authentication system 100 to perform individual voice authentication processing and voice analysis processing for multiple users, even if the audio signals are collected in an environment where multiple users speak simultaneously, such as a web conference.
[0022] An authentication analysis device P1, as an example of an authentication device and computer, is connected to the operator-side call terminal OP1, the registered speaker database DB, and the display DP to enable data transmission and reception. The authentication analysis device P1 may also be connected to the operator-side call terminal OP1, the registered speaker database DB, and the display DP via a network (not shown) to enable wired or wireless communication.
[0023] The authentication analysis device P1 acquires the voice signal of user US transmitted from the operator-side call terminal OP1, and analyzes the acquired voice signal, for example, frequency by frequency, to extract the individual speech features of user US. The authentication analysis device P1 refers to the registered speaker database DB and compares the extracted speech features with the speech features of multiple users previously registered in the registered speaker database DB to perform voice authentication of user US. The authentication analysis device P1 generates an authentication result screen SC that includes the authentication result of user US and sends it to the display DP for output. It goes without saying that the authentication result screen SC shown in Figure 1 is just an example and is not limited to it. The authentication result screen SC shown in Figure 1 includes, for example, the message "Matched the voice of Taro Yamada," which is the authentication result of user US.
[0024] As an example of a database, the registered speaker database DB is a so-called storage system, configured using storage media such as flash memory, HDD (Hard Disk Drive), or SSD (Solid State Drive). The registered speaker database DB stores (registers) the user information and speech features of multiple users in association with each user. The user information referred to here is information about the user, such as the username, user ID (Identification), or identification information assigned to each user. The registered speaker database DB may be configured integrally with the authentication analysis device P1.
[0025] The display DP is configured using, for example, an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display, and displays the authentication result screen SC transmitted from the authentication analysis device P1. The display DP may be configured integrally with the authentication analysis device P1.
[0026] In the example shown in Figure 1, the user-side call terminal UP1 collects the user US's spoken voice COM12 "This is Yamada Taro" and spoken voice COM14 "This is 123245678", converts them into voice signals, and transmits them to the operator-side call terminal OP1. The operator-side call terminal OP1 transmits the voice signals based on the user US's spoken voice COM12 and COM14 transmitted from the user-side call terminal UP1 to the authentication analysis device P1.
[0027] Furthermore, when the operator-side call terminal OP1 acquires audio signals that include the operator OP's spoken voice COM11 "Please tell me your name" and spoken voice COM13 "Please tell me your membership number," as well as the user US's spoken voice COM12 and spoken voice COM14, it separates and removes the audio signals based on the operator OP's spoken voice COM11 and COM13, extracts only the audio signals based on the user US's spoken voice COM12 and COM14, and transmits them to the authentication analysis device P1. As a result, the authentication analysis device P1 can improve the accuracy of user authentication by using only the voice signals of the person being authenticated.
[0028] Next, with reference to Figure 2, an example of the internal configuration of the authentication analysis device according to this embodiment will be described. Figure 2 is a block diagram showing an example of the internal configuration of the authentication analysis device according to this embodiment. The authentication analysis device P1 is composed of at least a communication unit 20, a processor 21, and a memory 22.
[0029] The communication unit 20 is connected to the operator-side call terminal OP1 and the registered speaker database DB to enable data communication. The communication unit 20 outputs the voice signal transmitted from the operator-side call terminal OP1 to the processor 21.
[0030] The processor 21 is composed of a semiconductor chip on which at least one of the following electronic devices is implemented: a CPU (Central Processing Unit), a DSP (Digital Signal Processor), a GPU (Graphical Processing Unit), or an FPGA (Field Programmable Gate Array). The processor 21 functions as a controller that oversees the overall operation of the authentication analysis device P1, performing control processing to coordinate the operation of each part of the authentication analysis device P1, data input / output processing between each part of the authentication analysis device P1, data calculation processing, and data storage processing.
[0031] The processor 21 uses programs and data stored in the ROM (Read Only Memory) 22A of the memory 22 to implement the functions of the speech segment detection unit 21A, the registration quality determination unit 21B, the feature extraction unit 21C, the comparison target setting unit 21D, the similarity calculation unit 21E, the authentication condition setting unit 21F, the sound collection condition measurement unit 21G during authentication, and the operation restriction setting unit 21H. During operation, the processor 21 uses the RAM (Random Access Memory) 22B of the memory 22 to temporarily store data or information generated or acquired by the processor 21 and each unit in the RAM 22B of the memory 22.
[0032] As an example of a detection unit, the speech segment detection unit 21A acquires the speech signal of the spoken voice during authentication (hereinafter referred to as the "speech voice signal"), analyzes the acquired speech voice signal, and detects the speech segment spoken by the user US (hereinafter referred to as the "first speech segment"). The speech segment detection unit 21A outputs the speech voice signal corresponding to at least one first speech segment detected from the speech voice signal (hereinafter referred to as the "first voice signal") to the feature extraction unit 21C. The speech segment detection unit 21A may also temporarily store the first voice signal of at least one first speech segment in the RAM 22B of the memory 22. If the speech segment detection unit 21A detects multiple first speech segments, it may concatenate the first voice signals of each detected first speech segment and output them to the feature extraction unit 21C. Furthermore, when registering the speech audio signal used for user US authentication, the speech segment detection unit 21A detects the speech segments of the audio data acquired from user US (hereinafter referred to as the second speech segment). The speech segment detection unit 21A outputs the speech audio signal corresponding to the second speech segment (hereinafter referred to as the second audio signal) to the registration quality determination unit 21B. If there are multiple second speech segments, the speech segment detection unit 21A may concatenate the second audio signals of each detected second speech segment and output them to the registration quality determination unit 21B.
[0033] As an example of a processing unit, the registration quality determination unit 21B acquires a second speech signal from the speech segment detection unit 21A, which is a combination of a second speech segment or multiple second speech segments. The registration quality determination unit 21B determines the quality of the acquired second speech signal. Quality is an indicator that shows the quality of the user's surrounding environment at the time of registration, the accuracy of the user's speech, or both, when the second speech signal is registered in the registered speaker database DB for each user prior to actual authentication (at the time of registration). In this embodiment, the authentication conditions imposed on the user at the time of actual authentication (see below) are determined based on this quality at the time of registration. The registration quality determination unit 21B determines the quality based, for example, on the length of the speech in the second speech signal (hereinafter referred to as speech length) or the number of sounds contained in the second speech signal. Note that the elements used by the registration quality determination unit 21B to determine the quality are not limited to speech length and number of sounds, but may also be the number of phonemes or the number of words. The registration quality determination unit 21B outputs the determined quality information to the feature extraction unit 21C or the authentication condition setting unit 21F.
[0034] As an example of a processing unit, the feature extraction unit 21C analyzes the characteristics of an individual's voice, for example by frequency, using one or more speech audio signals extracted by the speech segment detection unit 21A, and extracts speech features. The feature extraction unit 21C extracts speech features from the first audio signal of the first speech segment output from the speech segment detection unit 21A. The feature extraction unit 21C also extracts speech features from the second audio signal of the second speech segment output from the speech segment detection unit 21A. Note that the speech features of the second audio signal of the second speech segment may be registered in advance in the registered speaker database DB. The feature extraction unit 21C associates the extracted speech features of the first speech segment with the first audio signal from which these speech features were extracted, and outputs them to the similarity calculation unit 21E, or to the comparison target setting unit 21D, or temporarily stores them in the RAM 22B of the memory 22. The feature extraction unit 21C outputs the speech features of the second utterance section to the similarity calculation unit 21E by associating them with the second speech signal from which these speech features have been extracted, or by linking the speech features of the second utterance section with quality-related information obtained from the registration quality determination unit 21B and temporarily storing it in the RAM 22B of the memory 22.
[0035] The feature extraction unit 21C performs speech recognition on the content of the spoken audio signal. The method for speech recognition of the content of the spoken audio signal can be implemented using publicly known techniques. For example, it may be calculated as linguistic information by performing phoneme analysis on the spoken audio signal, or it may be implemented using other analysis methods.
[0036] As an example of a setting unit, the comparison target setting unit 21D obtains data of the speaker, user US, from the registered speaker database DB. Here, user US data refers to, for example, personal information such as user US's date of birth, name, or gender, or at least one of the voice data or feature quantities of the voice data related to utterances previously registered by user US. To set the speaker as user US, the comparison target setting unit 21D may, for example, identify the speaker as user US using the extracted speaker features output from the feature quantity extraction unit 21C, or it may identify the speaker as user US from the content entered by the speaker into the user-side call terminal UP1 (for example, name or ID). The comparison target setting unit 21D outputs the acquired user US data to the utterance interval detection unit 21A or the similarity calculation unit 21E.
[0037] As an example of an authentication unit, the similarity calculation unit 21E acquires speech features from the speech audio signal output by the feature extraction unit 21C. The similarity calculation unit 21E calculates the similarity between the speech features of a first speech segment and the speech features of a second speech segment, both acquired from the feature extraction unit 21C. Based on the calculated similarity, the similarity calculation unit 21E identifies the user corresponding to the speech audio signal (i.e., the audio signal transmitted from the user-side call terminal UP1) and performs user authentication to verify the user's identity.
[0038] The authentication condition setting unit 21F, as an example of a determination unit, sets authentication conditions based on quality information obtained from the registered quality determination unit 21B. Authentication conditions include, for example, the length of speech uttered by the user US, the content of the speech, or thresholds related to the determination. However, the authentication conditions are not limited to these.
[0039] As an example of a measurement unit, the authentication sound collection condition measurement unit 21G measures the sound collection conditions during authentication. Sound collection conditions include, for example, the noise, volume, degree of reverberation of the speech audio signal collected during authentication, or the number of phonemes contained in the speech audio signal. However, the sound collection conditions are not limited to these. The authentication sound collection condition measurement unit 21G outputs the measured sound collection conditions to the authentication condition setting unit 21F.
[0040] As an example of a setting unit, the operation restriction setting unit 21H sets restrictions on the actions that the user US can perform based on the quality of the spoken voice signal in the second utterance section. For example, if the authentication system 100 is installed in an ATM (Automatic Teller Machine), the operation restriction setting unit 21H will restrict actions such as money transfers or remittances if the quality of the spoken voice signal is poor. Note that the machine in which the authentication system 100 is installed is not limited to ATMs.
[0041] Based on these, the processor 21 sets authentication conditions for user identity verification based on the quality of the second voice signal determined by the registration quality determination unit 21B. The processor 21 acquires the user's utterance voice signal based on the set authentication conditions. The processor 21 authenticates whether the speaker is the person in question based on a comparison between the first voice signal of the first utterance section detected by the utterance section detection unit 21A and the second voice signal of the second utterance section.
[0042] Memory 22 includes, for example, ROM 22A which stores a program that defines various processes performed by the processor 21 and the data used during the execution of that program, and RAM 22B which serves as work memory used when executing various processes performed by the processor 21. ROM 22A contains a program that defines various processes performed by the processor 21 and the data used during the execution of that program. RAM 22B temporarily stores data or information generated or acquired by the processor 21 (for example, speech signals or speech feature quantities corresponding to each speech signal).
[0043] The display interface 23 connects the processor 21 and the display DP to enable data communication and outputs the authentication result screen SC generated by the similarity calculation unit 21E of the processor 21 to the display DP. Based on the authentication result of the processor 21, the display interface 23 displays the authentication status on the display DP, indicating whether or not the speaker is the person in question.
[0044] Next, the registration process for the speech voice signal for registration will be explained with reference to Figure 3. Figure 3 is a flowchart relating to the registration process for the speech voice signal for registration. Each process in the flowchart of Figure 3 is executed by the processor 21.
[0045] The flowchart in Figure 3 illustrates the process related to registration, that is, the registration of spoken audio signals that are stored in the registered speaker database DB beforehand.
[0046] The processor 21 begins receiving a speech signal for registration from the speaker (hereinafter referred to as the registration speech signal) (St10). In other words, in step St10, the speaker begins speaking to the user-side call terminal UP1.
[0047] Processor 21 terminates receiving the registered voice signal from the speaker (St11). In other words, in step St11, the speaker ends speaking to the user-side call terminal UP1.
[0048] The speech segment detection unit 21A detects a second speech segment of the registered voice signal acquired in the processing from step St10 to step St11 (St12).
[0049] The registration quality determination unit 21B determines the quality of the second voice signal of the second utterance section detected in the processing of step St12 (St13).
[0050] The registration quality determination unit 21B determines whether or not to reacquire the registered voice signal based on the quality determined in step St13 (St14). The registration quality determination unit 21B determines not to reacquire the signal if the quality is above a predetermined minimum required value, and determines to reacquire the signal if the quality is below a predetermined minimum required value. For example, if the speaker has not spoken a word, the speech length is 1 second, or the number of sounds is 1, the registration quality determination unit 21B determines to reacquire the registered voice signal. Note that the examples of cases in which the registration quality determination unit 21B determines to reacquire the signal are just examples and are not limited to these. Also, the processing in step St14 may be omitted from the flowchart shown in Figure 3.
[0051] If the registration quality determination unit 21B determines that the registered audio signal should be reacquired (St14, YES), the processor 21 returns to the process of step St10.
[0052] If the registration quality determination unit 21B determines that the registered speech signal should not be reacquired (St14, NO), the feature extraction unit 21C extracts speech features from the speech speech signal of the second speech interval (St15).
[0053] The feature extraction unit 21C links the quality determined in step St13 with the speech features extracted in step St15 and stores them in the registered speaker database DB (St16).
[0054] Next, with reference to Figure 4, an example of setting authentication conditions based on utterance length will be explained. Figure 4 is a diagram showing an example of setting authentication conditions based on utterance length.
[0055] Registered voice signals US10, US11, and US12 are registered voice signals of user US that were registered in the registered speaker database DB in the process shown in Figure 3.
[0056] In the example shown in Figure 4, the quality is "low" if the utterance length is less than 10 seconds, and "high" if the utterance length is 10 seconds or more. The threshold number of seconds for the quality to be "low" or "high" is just an example and is not limited to any specific range. Furthermore, the quality is not limited to just two levels, "low" and "high," but may be set to three levels, such as "low," "medium," and "high," or even four or more levels.
[0057] The authentication condition setting unit 21F changes the requested time based on the quality result. In the example shown in Figure 4, if the quality is "low," the requested time is 15 seconds, and if the quality is "high," the requested time is 7 seconds. The requested time is the total duration of utterances that the authentication system 100 requests from the speaker when performing authentication. Note that the length of the requested time is just an example and is not limited to these values. In the example shown in Figure 4, the judgment threshold is set to 70 regardless of the quality result. The judgment threshold is the threshold used by the similarity calculation unit 21E to determine the similarity between the speech features of the first utterance interval and the speech features of the second utterance interval. A higher judgment threshold requires a higher degree of similarity. Note that the value of the judgment threshold is just an example and is not limited to 70.
[0058] The content of the registered voice signal US10 is "a ka sa ta na de su" and the length of the utterance is 5 seconds. Since the length of the registered voice signal US10 is 5 seconds and less than 10 seconds, the quality is "low". As a result, the required time for the authentication conditions related to the registered voice signal US10 is 15 seconds, and the judgment threshold is 70.
[0059] The content of the registered voice signal US11 is "a ka sa ta na de su ha ma ya ra wa de su" and the length of the speech is 8 seconds. Since the length of the registered voice signal US11 is 8 seconds, which is less than 10 seconds, the quality is "low". As a result, the required time for authentication for the registered voice signal US11 is 15 seconds, and the judgment threshold is 70.
[0060] The content of the registered voice signal US12 is "a ka sa ta na de su i chi ni sa n shi go roku na na de su" and the length of the speech is 13 seconds. Since the length of the registered voice signal US12 is 13 seconds, which is more than 10 seconds, the quality is rated as "high". As a result, the required time for authentication for the registered voice signal US12 is 7 seconds, and the judgment threshold is 70.
[0061] Next, with reference to Figure 5, an example of setting authentication conditions based on utterance length and number of syllables will be explained. Figure 5 is a diagram showing an example of setting authentication conditions based on utterance length and number of syllables.
[0062] In the example shown in Figure 5, the quality is "low" if the utterance length is less than 10 seconds, and "high" if the utterance length is 10 seconds or more. The quality is also "low" if the number of syllables is less than 13, and "high" if the number of syllables is 13 or more. Note that the threshold time in seconds and the number of syllables for the quality to be "low" or "high" are examples only and are not limited. Furthermore, the quality is not limited to "low" and "high," but may be set to three levels such as "low," "medium," and "high," or even four or more levels.
[0063] Based on the quality results of the speech length and number of tones, the authentication condition setting unit 21F changes the requested time. The lower of the two quality levels, speech length and number of tones, is used as the quality of the registered speech signal. In the example shown in Figure 5, if the quality is "low," the requested time is 15 seconds, and if the quality is "high," the requested time is 7 seconds. Note that the length of the requested time is just an example and is not limited to these. In the example shown in Figure 5, the judgment threshold is set to 70 regardless of the quality result.
[0064] The content of the registered voice signal US10 is "a ka sa ta na de su" and has 7 syllables. The speech length of the registered voice signal US10 is 5 seconds, which is less than 10 seconds, so the quality related to speech length is "low". The number of syllables is 7, which is less than 13, so the quality related to the number of syllables is "low". Since both the quality of speech length and the number of syllables are "low", the quality of the registered voice signal US10 is "low". As a result, the required time for authentication conditions related to the registered voice signal US10 is 15 seconds, and the judgment threshold is 70.
[0065] The content of the registered voice signal US11 is "a ka sa ta na de su ha ma ya ra wa de su", and it has 14 syllables. The speech length of the registered voice signal US11 is 8 seconds, which is less than 10 seconds, so the quality related to speech length is "low". The number of syllables is 14, which is 13 or more, so the quality related to the number of syllables is "high". The quality of the number of syllables is "high", but the quality of the speech length is "low", so the quality of the registered voice signal US11 is "low". As a result, the required time for authentication conditions related to the registered voice signal US11 is 15 seconds, and the judgment threshold is 70.
[0066] The content of the registered voice signal US12 is "a ka sa ta na de su i chi ni sa n shi go ro ku na na de su" and has 20 syllables. The speech length of the registered voice signal US12 is 13 seconds, which is more than 10 seconds, so the quality related to speech length is "high". The number of syllables is 20, which is more than 13 syllables, so the quality related to the number of syllables is "high". Since both the quality of speech length and the number of syllables are "high", the quality of the registered voice signal US12 is "high". As a result, the required time for authentication conditions related to the registered voice signal US12 is 7 seconds, and the judgment threshold is 70.
[0067] Next, with reference to Figure 6, an example of setting utterance content as an authentication condition according to the quality of the registered voice signal will be explained. Figure 6 is a diagram showing an example of setting utterance content as an authentication condition according to the quality of the registered voice signal. The method for determining the quality in Figure 6 is the same as the method for determining it in Figure 4.
[0068] The authentication condition setting unit 21F specifies a phrase to prompt the user US to speak based on the quality result. In the example shown in Figure 6, if the quality is "low," the feature extraction unit 21C performs speech recognition on the registered voice signal, analyzes the utterance content, and outputs it to the authentication condition setting unit 21F. The authentication condition setting unit 21F determines the phrase to specify to the user US based on the speech recognition result obtained from the feature extraction unit 21C. If the quality is "low" and the authentication condition setting unit 21F specifies the utterance content, no request time is specified. If the quality is "high," no phrase is specified to the user US. However, even if the quality is "high," a phrase may be specified to the user US, and a shorter sentence than when the quality is "low" may be specified to improve user convenience. In the example shown in Figure 6, if the quality is "high," the request time is set to 7 seconds as an authentication condition. However, the request time when the quality is "high" is just an example and is not limited to 7 seconds. In the example shown in Figure 6, the judgment threshold is a constant value of 70 regardless of the quality. Furthermore, the judgment threshold does not have to be a fixed value and may be changed depending on the quality.
[0069] The content of the registered voice signal US10 is "akasata na desu". Since the quality of the registered voice signal US10 is "low", the feature extraction unit 21C performs speech recognition on the registered voice signal US10, and the recognition result is "akasata na desu". Based on the recognition result obtained from the feature extraction unit 21C, the authentication condition setting unit 21F specifies the phrase as "akasata na desu". The authentication condition setting unit 21F sets the request time to "unspecified" and the judgment threshold to 70.
[0070] The content of the registered voice signal US12 is "akasata na desu ichi ni san shi goro ku nana desu". Since the quality of the registered voice signal US12 is "high", the feature extraction unit 21C does not perform speech recognition. The authentication condition setting unit 21F sets the request time to 7 seconds and the judgment threshold to 70 as authentication conditions. The authentication condition setting unit 21F does not specify any words because the quality of the registered voice signal US12 is "high".
[0071] As a result, the authentication system 100 can authenticate quickly while maintaining high authentication accuracy by specifying and having the user US speak a phrase based on the content of the registered voice signal, depending on the quality of the registered voice signal. Furthermore, if the quality of the registered voice signal is high, the authentication system 100 can reduce the effort required of the user US by not specifying the content of the speech.
[0072] Next, with reference to Figure 7, we will explain an example in which an operator performs identity verification authentication based on the identity verification text displayed on the screen. Figure 7 is a diagram showing an example in which an operator performs identity verification authentication based on the authentication text displayed on the screen.
[0073] If the quality shown in Figure 6 is "low," the authentication system 100 uses speech recognition to specify the words to be spoken by the user US. Figure 7 illustrates an example in which the words to be spoken by the user US specified by the method in Figure 6 are displayed on a screen that the operator sees during user US identity verification authentication (hereinafter referred to as the operator screen), and the operator OP has the user US speak the specified words.
[0074] First, let's explain the example shown in Case CC. Screen SC1 is an example of an operator screen displayed on Display DP.
[0075] Frame FR1 displays the speaker's registration information. Frame FR1 displays the following information as the speaker's registration information: "Caller ID," "Registered Name," "Registered Address," "Age," and "Speaker Registration Status." "Caller ID" is, for example, a telephone number. "Speaker Registration Status" indicates whether or not a registered voice signal is stored in the registered speaker database. If a registered voice signal is stored in the registered speaker database, the quality associated with the registered voice signal is also displayed. For example, frame FR1 displays "Caller ID" as xx-xxxx-xxxx, "Registered Name" as A-ta A-o, "Registered Address" as ABCDEFG, "Age" as 33, and "Speaker Registration Status" as Yes (Quality: Low).
[0076] In box FR2, the candidates who are considered to be the speaker are displayed as the speaker authentication result. Next to each candidate, the probability that the speaker is one of the candidates is displayed. In the example in box FR2, the probability is shown as a percentage, but it is not limited to this and can also be displayed as "low, medium, high". In box FR2, the authentication results are displayed as "A. A.: 70%", "B. B.: 25%", and "C. C.: 5%".
[0077] Frame FR3 displays the text to be spoken by the speaker (hereinafter referred to as the "authentication text"). Frame FR3 displays "Hamayarawa desu hachikyujuuzero desu" as the authentication text.
[0078] Button BT1 is used to start or stop authentication.
[0079] Operator OP speaks OP10, "Please say 'Hamayarawa, it's 890zero'," based on the authentication text displayed in frame FR3 of screen SC1. User US speaks US13, "Hamayarawa, it's 890zero," based on Operator OP's speech.
[0080] Next, we will explain the example shown in Case CD. Screen SC2 is an example of the operator screen displayed on Display DP.
[0081] Frame FR5 displays the speaker's registration information. Frame FR5 displays the "Caller ID," "Registered Name," "Registered Address," "Age," and "Speaker Registration Status." For example, frame FR2 displays "Caller ID" as xx-xxxx-xxxx, "Registered Name" as B-yama B-ro, "Registered Address" as GFEDCBA, "Age" as 44, and "Speaker Registration Status" as Yes (Quality: High).
[0082] In frame FR6, the speaker authentication results show the candidates who are considered to be the speaker. Next to each candidate, the probability that the speaker is one of the candidates is displayed. In frame FR2, the authentication results are displayed as "A. Tadashi: 15%", "B. Yamada: 60%", and "C. Kawashima: 25%".
[0083] Frame FR7 displays "Hamayarawa desu" as the authentication text. Case CD is of "high" quality, so a shorter text is specified as the authentication text than Case CC, which is of "low" quality.
[0084] Button BT2 is used to start or stop authentication.
[0085] Operator OP speaks OP11, "Please say 'Hamayarawa'," based on the authentication text displayed in frame FR7 of screen SC2. User US speaks US14, "Hamayarawa," based on Operator OP's speech.
[0086] In Figure 7, the operator OP reads aloud the authentication text displayed on the operator screen and has the user US speak it. However, the authentication text may also be played via automated voice and spoken by the user US.
[0087] This allows the authentication system 100 to perform authentication without the user US and operator OP having to worry about the length of their speech. Furthermore, if the quality of the registered voice signal is high, the authentication system 100 can shorten the time required for authentication by specifying a shorter phrase to the user US. Also, if the quality of the registered voice signal is low, the authentication system 100 can maintain high authentication accuracy by specifying a longer phrase to the user US than when the quality is high, thereby preventing authentication failures or retrying.
[0088] Next, with reference to Figure 8, we will explain an example of performing identity verification authentication based on identity verification text displayed on the user's phone terminal. Figure 8 is a diagram showing an example of performing identity verification authentication based on identity verification text displayed on the user's phone terminal.
[0089] First, let's explain the example shown in Case CE. Case CE is an example of a situation where the quality of the registered voice signal is low. Screen SC3 is an example of a screen displayed on the user's call terminal UP1.
[0090] Screen SC3 displays the message, "Please say the identity verification phrase 'hamayarawa desu hachikyujuzero desu'." Frame FR9 displays the identity verification phrase "hamayarawa desu hachikyujuzero desu," but the displayed text differs depending on the user US.
[0091] User US sees the content displayed on screen SC3 and speaks voice US13, "Hamayarawa desu, hachikyujuzero desu."
[0092] Next, we will explain the example shown in Case CF. Case CF is an example where the quality of the registered voice signal is high. Screen SC4 is an example of a screen displayed on the user-side call terminal UP1.
[0093] Screen SC4 displays "Please say the identity verification phrase 'hamayarawa desu'." Frame FR9 displays "hamayarawa desu" as the identity verification phrase.
[0094] User US sees the content displayed on screen SC4 and speaks the voice US13 "Hamayarawa desu".
[0095] This allows the authentication system 100 to perform authentication without requiring the user US to be concerned about the length of their speech. Furthermore, this enables the authentication system 100 to perform user US authentication unattended, without the need for an operator OP or other human intervention.
[0096] Next, with reference to Figures 9 and 10, we will explain an example of resetting the authentication conditions based on the measurement results of the sound collection conditions during authentication after the authentication conditions have been set. Figure 9 shows an example of resetting the required time for the authentication conditions based on the measurement results of the sound collection conditions during authentication after the authentication conditions have been set. Figure 10 shows an example of resetting the threshold for the authentication conditions based on the measurement results of the sound collection conditions during authentication after the authentication conditions have been set.
[0097] The authentication sound collection condition measurement unit 21G measures the noise, volume, degree of reverberation of the speech audio signal collected during authentication, or the number of phonemes contained in the speech audio signal, as authentication sound collection conditions (hereinafter referred to as authentication sound collection conditions). Figures 9 and 10 show an example in which the authentication conditions are reset based on the authentication sound collection conditions measured after the authentication conditions have been set once (hereinafter referred to as initial authentication conditions).
[0098] If the noise in the spoken audio signal exceeds a predetermined noise level, i.e., if there is a lot of noise, the authentication condition setting unit 21F increases the request time by 3 seconds as an authentication condition.
[0099] If the volume of the spoken audio signal is below a predetermined value related to volume, that is, if the volume is low, the authentication condition setting unit 21F increases the request time by 3 seconds as an authentication condition.
[0100] If the number of phonemes in the spoken audio signal is less than a predetermined value related to the number of phonemes, that is, if the number of phonemes is small, the authentication condition setting unit 21F increases the request time by 3 seconds as an authentication condition.
[0101] If the reverberation of the spoken audio signal exceeds a predetermined value related to reverberation, that is, if the reverberation is high, the authentication condition setting unit 21F increases the request time by 5 seconds as an authentication condition.
[0102] The requested durations for noise, volume, phoneme count, and reverberation are examples only and not limited to these.
[0103] If the quality of the spoken audio signal is "low," the initial authentication conditions are a request time of 15 seconds and a judgment threshold of 70. Note that the initial authentication conditions are just examples and are not limited to these. If there is a lot of noise as an authentication sound collection condition, the authentication condition setting unit 21F increases the request time by 3 seconds. As a result, the authentication conditions after resetting are a request time of 18 seconds and a judgment threshold of 70. If the volume is low as an authentication sound collection condition, the authentication condition setting unit 21F increases the request time by 3 seconds. As a result, the authentication conditions after resetting are a request time of 18 seconds and a judgment threshold of 70.
[0104] If the quality of the spoken audio signal is "high," the initial authentication conditions are a request time of 7 seconds and a judgment threshold of 70. Note that the initial authentication conditions are just an example and are not limited to these. If the sound collection conditions during authentication are low volume and noisy, the authentication condition setting unit 21F increases the request time by a total of 6 seconds. As a result, the authentication conditions after resetting are a request time of 13 seconds and a judgment threshold of 70. If the sound collection conditions during authentication are low in phonemes, the authentication condition setting unit 21F increases the request time by 3 seconds. As a result, the authentication conditions after resetting are a request time of 10 seconds and a judgment threshold of 70. If the sound collection conditions during authentication are high in reverberation, the authentication condition setting unit 21F increases the request time by 5 seconds. As a result, the authentication conditions after resetting are a request time of 12 seconds and a judgment threshold of 70. If the sound collection conditions during authentication are good, the authentication conditions are the same as the initial authentication conditions.
[0105] Next, referring to Figure 10, we will explain an example of resetting the authentication threshold based on the measurement results of the sound collection conditions during authentication after the authentication conditions have been set.
[0106] If the noise in the spoken audio signal exceeds a predetermined noise level, i.e., if there is a lot of noise, the authentication condition setting unit 21F lowers the judgment threshold by 10 as an authentication condition.
[0107] If the volume of the spoken audio signal is below a predetermined value related to volume, that is, if the volume is low, the authentication condition setting unit 21F lowers the judgment threshold by 15 as an authentication condition.
[0108] If the number of phonemes in the spoken audio signal is less than a predetermined value related to the number of phonemes, that is, if the number of phonemes is small, the authentication condition setting unit 21F lowers the judgment threshold by 10 as an authentication condition.
[0109] If the reverberation of the spoken audio signal exceeds a predetermined value related to reverberation, that is, if the reverberation is high, the authentication condition setting unit 21F lowers the judgment threshold by 20 as an authentication condition.
[0110] Note that the threshold values for lowering noise, volume, phoneme count, and reverberation are examples only and are not limited to these.
[0111] If the quality of the spoken audio signal is "low," the initial authentication conditions are a request time of 15 seconds and a judgment threshold of 70. Note that the initial authentication conditions are just examples and are not limited to these. If there is a lot of noise as an authentication sound collection condition, the authentication condition setting unit 21F lowers the judgment threshold by 10. As a result, the authentication conditions after resetting are a request time of 15 seconds and a judgment threshold of 60. If the volume is low as an authentication sound collection condition, the authentication condition setting unit 21F lowers the judgment threshold by 15. As a result, the authentication conditions after resetting are a request time of 15 seconds and a judgment threshold of 55.
[0112] If the quality of the spoken audio signal is "high," the initial authentication conditions are a request time of 7 seconds and a judgment threshold of 70. Note that the initial authentication conditions are just an example and are not limited to these. If the sound collection conditions during authentication are low volume and noisy, the authentication condition setting unit 21F lowers the judgment threshold by a total of 25. As a result, the authentication conditions after resetting are a request time of 7 seconds and a judgment threshold of 45. If the sound collection conditions during authentication are low in phonemes, the authentication condition setting unit 21F lowers the judgment threshold by 10. As a result, the authentication conditions after resetting are a request time of 7 seconds and a judgment threshold of 60. If the sound collection conditions during authentication are high in reverberation, the authentication condition setting unit 21F lowers the judgment threshold by 20. As a result, the authentication conditions after resetting are a request time of 7 seconds and a judgment threshold of 50. If the sound collection conditions during authentication are good, the authentication conditions are the same as the initial authentication conditions.
[0113] As a result, the authentication system 100 can perform authentication with high accuracy regardless of the quality of the registered audio signal, by increasing the request time, lowering the judgment threshold, or both, if the sound collection conditions during authentication are poor.
[0114] Next, the process related to speaker authentication will be explained with reference to Figure 11. Figure 11 is a flowchart of the process related to speaker authentication. Each process in Figure 11 is executed by the processor 21.
[0115] The comparison target setting unit 21D sets the person to be used for authentication from among multiple people registered in the registered speaker database DB when performing authentication to verify the identity of the speaker who is the person to be authenticated (St20).
[0116] The comparison target setting unit 21D obtains information (St21) regarding the quality of the registered voice signal of the person to be compared, as set in the process of step St20, from the registered speaker database DB. The comparison target setting unit 21D outputs the obtained information to the authentication condition setting unit 21F. The comparison target setting unit 21D also obtains the registered feature quantities of the registered voice signal of the person to be compared from the registered speaker database DB and outputs them to the similarity calculation unit 21E.
[0117] The authentication sound collection condition measurement unit 21G measures the authentication sound collection conditions (St22). Note that the process in step St22 may be omitted from the flowchart shown in Figure 11.
[0118] The authentication condition setting unit 21F sets the authentication conditions (St23) based on the quality information obtained from the comparison target setting unit 21D in the process of step St21.
[0119] The processor 21 sends a signal to the communication unit 20 to initiate the authentication process (St24). The communication unit 20 sends an instruction to the operator-side call terminal OP1 to initiate authentication.
[0120] The authentication condition setting unit 21F obtains the registered feature quantities of the speaker's registered speech signal from the comparison target setting unit 21D. Based on the obtained registered feature quantities, the authentication condition setting unit 21F specifies the wording of the utterance to be used for authentication (St25). Note that the processing in step St25 may be omitted from the flowchart shown in Figure 11.
[0121] The processor 21 begins receiving the speech audio signal used for authentication (St26). The processor 21 outputs the acquired speech audio signal to the feature extraction unit 21C.
[0122] The authentication sound collection condition measurement unit 21G measures the authentication sound collection conditions (St27). The authentication sound collection condition measurement unit 21G outputs the measured authentication sound collection condition information to the authentication condition setting unit 21F. Note that the processing in step St27 may be omitted from the flowchart shown in Figure 11.
[0123] The authentication condition setting unit 21F resets the authentication conditions based on the authentication sound collection conditions obtained in step St27 (St28). The processor 21 acquires the spoken audio signal based on the authentication conditions reset by the authentication condition setting unit 21F in step St28. The processor 21 outputs the acquired spoken audio signal to the feature extraction unit 21C. Note that the processing in step St28 may be omitted from the flowchart shown in Figure 11.
[0124] The processor 21 sends a signal to the communication unit 20 to terminate the authentication process, that is, to terminate the reception of the speech voice signal used for authentication (St29). The communication unit 20 sends an instruction to the operator-side call terminal OP1 to terminate authentication.
[0125] The feature extraction unit 21C extracts speech features from the speech signal obtained in the processing of step St26 or step St28 (St30). The feature extraction unit 21C outputs the extracted speech features to the similarity calculation unit 21E.
[0126] The similarity calculation unit 21E calculates the similarity (St31) based on the registered features obtained in step St21 and the speech features obtained in the processing of step St30.
[0127] The similarity calculation unit 21E determines whether the similarity calculated in step St31 is equal to or greater than a predetermined threshold (St32). If the similarity calculation unit 21E determines that the similarity is equal to or greater than a predetermined threshold (St32, YES), it outputs a signal to the communication unit 20, the display I / F 23, or both, indicating that the speaker's identity verification has been successful.
[0128] If the similarity calculation unit 21E determines that the similarity is below a predetermined threshold (St32, NO), it determines whether or not to continue the authentication process (St34).
[0129] If the similarity calculation unit 21E determines to continue the authentication process (St34, YES), the processor 21 returns to the process in step St22. If the process in step St22 is omitted, the processor 21 returns to the process in step St23.
[0130] If the similarity calculation unit 21E determines that it will not continue the authentication process (St35, NO), it outputs a signal to the communication unit 20, the display I / F 23, or both indicating that the speaker's identity verification has failed.
[0131] Next, referring to Figure 12, we will explain an example of restricting operations after successful authentication based on the quality of the registered voice signal. Figure 12 is a diagram illustrating an example of restricting operations after successful authentication based on the quality of the registered voice signal.
[0132] The registered voice signals US10 and US12 shown in Figure 12 are the same as those shown in Figure 4. Therefore, the explanation of how the authentication conditions are set is omitted in Figure 12.
[0133] If the authentication system 100 is installed, for example, in a bank ATM, the operation restriction setting unit 21H may restrict operations (for example, deposits) after successful authentication of the speaker's identity based on the quality of the registered voice signal. Although the installation of the authentication system 100 is not limited to bank ATMs, for the sake of explanation, the authentication system 100 will be assumed to be installed in a bank ATM here.
[0134] Case CG is an example of authentication when the quality of the registered voice signal is low. In Case CG, because the quality of the registered voice signal is low, the operation restriction setting unit 21H restricts the operation mode and operates in restricted mode. In restricted mode, for example, only the inquiry of account balance and deposits are possible. Note that the operations possible in restricted mode are examples and are not limited to these.
[0135] Case CH is an example of authentication when the quality of the registered voice signal is high. In Case CH, because the quality of the registered voice signal is high, the operation restriction setting unit 21H operates in normal mode without restricting the operation mode. In normal mode, all operations are possible, such as checking the account balance, deposits, transfers, or remittances. Note that the operations possible in normal mode are examples and are not limited to these.
[0136] As a result, the authentication system 100 can reduce the risk of misidentification by imposing operational restrictions on the machine on which it is installed if the quality of the registered voice signal is poor.
[0137] Next, with reference to Figure 13, the process of setting operational restrictions based on the quality of the registered audio signal will be explained. Figure 13 is a flowchart of the process of setting operational restrictions based on the quality of the registered audio signal. Each process in the flowchart of Figure 13 is executed by the processor 21. Processes in the flowchart of Figure 13 that are the same as those in the flowchart of Figure 11 are denoted by the same symbols and their explanations are omitted.
[0138] If the authentication related to verifying the speaker's identity is successful in step St34, the operation restriction setting unit 21H determines whether the quality of the registered voice signal is high or not (St36).
[0139] If the operation restriction setting unit 21H determines that the quality of the registered audio signal is high (St36, YES), it sets the operation mode to normal mode (St37).
[0140] If the operation restriction setting unit 21H determines that the quality of the registered audio signal is low (St36, NO), it sets the operation mode to restriction mode (St38).
[0141] As described above, the authentication system according to this embodiment (for example, the authentication system 100) includes an acquisition unit (for example, a user-side call terminal UP1) that acquires the voice signal of the speaker's utterance. The authentication system includes a detection unit (for example, a voice segment detection unit 21A) that detects a first utterance segment in which the speaker is speaking from the acquired voice signal and a second utterance segment in which the speaker is speaking from the voice signals of a database in which the voice signals of multiple speakers are registered. The authentication system includes a determination unit (for example, an authentication condition setting unit 21F) that compares the first voice signal of the first utterance segment with the second voice signal of the second utterance segment and determines the authentication conditions for authentication using the first voice signal based on the length of the second voice signal of the second utterance segment or the number of sounds included in the second utterance segment. The authentication system includes an authentication unit (for example, a similarity calculation unit 21E) that performs authentication of the speaker based on the determined authentication conditions.
[0142] As a result, the authentication system according to this embodiment can determine the authentication conditions required of the user during authentication based on the length or number of sounds in the registered voice signal, and thus the authentication conditions can be changed for each user. This allows the authentication system to determine the speech duration during authentication according to the total length of the user's speech acquired during registration, thereby improving user convenience.
[0143] Furthermore, the authentication unit of the authentication system according to this embodiment starts authentication when the length of the second utterance is greater than or equal to a first predetermined value, and the length of the first utterance is greater than or equal to a second predetermined value. The authentication unit also starts authentication when the length of the second utterance is less than the first predetermined value, and the length of the first utterance is greater than or equal to a third predetermined value which is greater than the second predetermined value. As a result, the authentication system sets a number of seconds that is estimated to be sufficient for judgment at the time of authentication based on the quality of the utterance length of the registered voice signal and requests the user to utter, thereby preventing the user from uttering for more than necessary or from uttering for too short a time, which would cause authentication to fail. As a result, the authentication system 100 can determine the utterance time at the time of authentication according to the total length of the user's utterance acquired at the time of registration, thereby improving user convenience.
[0144] Furthermore, the authentication unit of the authentication system according to this embodiment starts authentication when the number of sounds included in the second utterance is equal to or greater than a fourth predetermined value and the length of the first utterance is equal to or greater than the second predetermined value. The authentication unit also starts authentication when the number of sounds included in the second utterance is less than the fourth predetermined value and the length of the first utterance is equal to or greater than a third predetermined value which is greater than the second predetermined value. As a result, the authentication system sets a number of seconds that is estimated to be sufficient for stability during authentication based on the quality of the number of sounds in the registered voice signal and requests the user to utter, thereby preventing the user from uttering for longer than necessary or failing authentication due to uttering for too short a time. This allows the authentication system to improve user convenience during authentication.
[0145] Furthermore, the authentication unit of the authentication system according to this embodiment starts authentication when the length of the second utterance is greater than or equal to a first predetermined value and the number of sounds included in the second utterance is greater than or equal to a fourth predetermined value, and the length of the first utterance is greater than or equal to the second predetermined value. The authentication unit also starts authentication when the length of the second utterance is less than the first predetermined value or the number of sounds included in the second utterance is less than the fourth predetermined value, and the length of the first utterance is greater than or equal to a third predetermined value which is greater than the second predetermined value. As a result, the authentication system can determine the utterance length required from the user at the time of authentication according to the utterance length and number of sounds of the registered voice signal. Since the authentication system can determine the authentication conditions from the utterance length and the number of sounds, it can improve user convenience and perform authentication with higher accuracy.
[0146] Furthermore, if the length of the second utterance segment is less than the first predetermined value, the determination unit of the authentication system according to this embodiment determines text to prompt the speaker to speak from the utterance content included in the audio signal of the second utterance segment. As a result, the authentication system can authenticate in a short time while maintaining high authentication accuracy by specifying and prompting the user to speak a phrase based on the utterance content of the registered audio signal according to the utterance length of the registered audio signal. In addition, if the utterance length of the registered audio signal is sufficiently long, the authentication system can save the user the trouble of specifying the utterance content.
[0147] Furthermore, the authentication system according to this embodiment further includes a first display unit (e.g., a display DP) that displays a screen for the operator to refer to during authentication, and the determination unit causes text to be displayed on the first display unit. This allows the authentication system to enable users and operators to perform authentication without worrying about the length of their speech. In addition, if the quality of the registered voice signal is high, the authentication system can shorten the time required for authentication by specifying a short phrase to the user. In addition, if the quality of the registered voice signal is low, the authentication system can maintain high authentication accuracy by specifying a longer phrase to the user than when the quality is high, thereby preventing authentication failures or retrying.
[0148] Furthermore, the authentication system according to this embodiment further includes a second display unit (for example, the user-side call terminal UP1) that displays a screen for the speaker to refer to during authentication, and the determination unit displays text on the second display unit. The authentication system can perform authentication without requiring the user to worry about the length of their speech. In addition, this allows the authentication system to perform user authentication unmanned without the need for an operator or other person to be involved.
[0149] Furthermore, the authentication system according to this embodiment further includes a measurement unit that measures at least one of the noise, volume, number of phonemes, or reverberation magnitude of the audio signal in the first utterance section, and a determination unit sets authentication conditions based on the measurement results obtained from the measurement unit. As a result, the authentication system can authenticate with high accuracy regardless of the quality of the registered audio signal by increasing the required time, lowering the judgment threshold, or both, if the sound collection conditions at the time of authentication are poor.
[0150] Furthermore, the authentication conditions of the authentication system according to this embodiment are the length of the voice signal or the threshold for determining the identity of the speaker. This allows the authentication system to specify the length of speech to be given to each user according to the quality of the registered voice signal of each user, thereby improving user convenience. In addition, the authentication system can change the threshold for determination according to the quality of the user's registered voice signal, enabling flexible authentication tailored to each user.
[0151] Furthermore, the authentication system according to this embodiment further includes a restriction setting unit (for example, an action restriction setting unit 21H) that restricts the actions that the speaker can perform after authentication if the length of the second utterance is less than a first predetermined value. This allows the authentication system to reduce the risk of misjudgment by imposing action restrictions on the machine on which the authentication system is installed when the quality of the registered voice signal is low.
[0152] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention. [Industrial applicability]
[0153] The technology disclosed herein is useful as an authentication system and authentication method that improves user convenience by determining the speech duration at the time of authentication based on the total length of the user's speech acquired at the time of registration. [Explanation of Symbols]
[0154] NW Network UP1 User-side calling terminal OP1 Operator-side call terminal US User OP Operator COM11, COM12, COM13, COM14 Speech Voice P1 Authentication Analysis Device DB Registered Speaker Database DP Display SC Authentication Results Screen 20 Communications Department 21 processors 21A Speech interval detection unit 21B Registration Quality Determination Unit 21C Feature Extraction Unit 21D Comparison Target Setting Section 21E Similarity calculation part 21F Authentication Condition Setting Section 21G Authentication Sound Collection Condition Measurement Unit 22H Operation Limit Setting Section 22 memory 22A ROM 22B RAM 23 Display I / F US10, US11, US12 Registered Audio Signals CA,CB,CC,CD,CE,CF,CG,CH Case SC1, SC2, SC3, SC4 screens FR1,FR2,FR3,FR4,FR5,FR6,FR7 frame BT1, BT2 buttons OP10, OP11, US13, US14 Speech Voice
Claims
1. An acquisition unit that acquires the audio signal of the speaker's utterance, A detection unit that detects a first speech segment spoken by the speaker from the acquired speech signal, and a second speech segment spoken by the speaker from the speech signal in a database in which the speech signals of multiple speakers are registered, A determination unit that compares the first voice signal of the first speech segment with the second voice signal of the second speech segment and determines the authentication conditions for using the first voice signal based on the length of the second voice signal of the second speech segment or the number of sounds included in the second speech segment, The system comprises an authentication unit that authenticates the speaker based on the determined authentication conditions, The authentication unit, The authentication is initiated when the length of the second utterance interval is equal to or greater than the first predetermined value, and the length of the first utterance interval is equal to or greater than the second predetermined value. The authentication is initiated when the length of the second utterance is less than the first predetermined value, and the length of the first utterance is greater than or equal to a third predetermined value which is greater than the second predetermined value. Authentication system.
2. The authentication unit, The authentication is initiated when the number of sounds included in the second speech segment is equal to or greater than the fourth predetermined value, and the length of the first speech segment is equal to or greater than the second predetermined value. Authentication is initiated when the number of sounds included in the second speech segment is less than the fourth predetermined value, and the length of the first speech segment is greater than or equal to the third predetermined value which is greater than the second predetermined value. The authentication system according to claim 1.
3. The authentication unit, The authentication is initiated when the length of the second speech segment is equal to or greater than the first predetermined value and the number of sounds included in the second speech segment is equal to or greater than the fourth predetermined value, and the length of the first speech segment is equal to or greater than the second predetermined value. Authentication is initiated when the length of the second speech segment is less than the first predetermined value or the number of sounds included in the second speech segment is less than the fourth predetermined value, and the length of the first speech segment is greater than or equal to a third predetermined value which is greater than the second predetermined value. The authentication system according to claim 1.
4. The aforementioned determination unit, If the length of the second utterance segment is less than a first predetermined value, text prompting the speaker to speak is determined from the utterance content included in the audio signal of the second utterance segment. The authentication system according to claim 1.
5. The system further includes a first display unit that displays a screen that the operator refers to during the authentication process, The determination unit causes the text to be displayed on the first display unit. The authentication system according to claim 4.
6. The system further includes a second display unit that displays a screen that the speaker refers to during authentication. The determination unit causes the text to be displayed on the second display unit. The authentication system according to claim 4.
7. The system further includes a measuring unit that measures at least one of the noise, volume, number of phonemes, or reverberation magnitude of the speech signal in the first speech segment, The determination unit sets the authentication conditions based on the measurement results obtained from the measurement unit. The authentication system according to claim 4.
8. The authentication condition is the length of the audio signal or the threshold for determining the identity of the speaker. The authentication system according to claim 7.
9. If the length of the second utterance interval is less than a first predetermined value, the system further includes a restriction setting unit that imposes restrictions on the actions that the speaker can perform after authentication. The authentication system according to claim 1.
10. An authentication method performed by one or more computers, The audio signal of the speaker's utterance is acquired, From the acquired audio signal, a first utterance segment spoken by the speaker is detected, and from the audio signals in the database in which the audio signals of multiple speakers are registered, a second utterance segment spoken by the speaker is detected. The first audio signal of the first speech segment and the second audio signal of the second speech segment are compared, The authentication conditions for authentication using the first voice signal are determined based on the length of the second voice signal in the second speech segment or the number of sounds included in the second speech segment. Based on the determined authentication conditions, the speaker is authenticated. The authentication is initiated when the length of the second utterance interval is equal to or greater than the first predetermined value, and the length of the first utterance interval is equal to or greater than the second predetermined value. The authentication is initiated when the length of the second utterance is less than the first predetermined value, and the length of the first utterance is greater than or equal to a third predetermined value which is greater than the second predetermined value. Authentication method.
Citation Information
Patent Citations
Communication device, method and program for registering voice print
JP2016053598A
Registered utterance division device, speaker likelihood evaluation device, speaker identification device, registered utterance division method, speaker likelihood evaluation method, and program
JP2017187642A
Voice identification enrollment
US20190341055A1
Speaker Identification with Ultra-Short Speech Segments for Far and Near Field Voice Assistance Applications
US20200152206A1
User authentication method and apparatus
US20200175993A1