Authentication device and authentication method

JP7912271B2Active Publication Date: 2026-08-28PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023549434
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-09-27
Filing Date
2022-08-29
Publication Date
2026-08-28
Estimated Expiration
2042-08-29

AI Technical Summary

Benefits of technology

【0008】 本開示によれば、発話音声を用いた話者の音声認証精度を向上できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007912271000001
    Figure 0007912271000001
  • Figure 0007912271000002
    Figure 0007912271000002
  • Figure 0007912271000003
    Figure 0007912271000003
Patent Text Reader

Abstract

An authentication device according to the present invention comprises: an acquisition unit that acquires a voice signal of a speaker; a detection unit that detects a first speech segment spoken by the speaker; and an authentication unit that authenticates the speaker on the basis of a comparison of the voice signal of the first speech segment and a database. If it has been determined that the speaker cannot be authenticated, the detection unit detects a second speech segment differing from the first speech segment, and the authentication unit authenticates the speaker on the basis of a comparison of voice signals of the first speech segment and second speech segment and the database.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an authentication device and an authentication method. [Background Art]

[0002] Patent Document 1 discloses an authentication device for verifying the identity of a speaker who makes a call using a telephone terminal connected to a telephone network, the authentication device determining the identity of the speaker based on a voice recognition authentication result. The authentication device stores predetermined voiceprint information, a first keyword, and a second keyword, acquires voiceprint information from voice received by a receiving means, and performs voiceprint authentication by collating the acquired voiceprint information with the stored predetermined voiceprint information. The authentication device transmits a voice message prompting the speaker to utter the first keyword to the telephone terminal, and then determines whether the content of the speaker's voice received by the receiving means corresponds to the first keyword stored in a storage means. When the authentication result obtained using the voiceprint information differs from the voice recognition authentication result obtained using the first keyword, the authentication device transmits a voice message prompting the speaker to utter the second keyword to the telephone terminal, and then determines whether the content of the speaker's voice received by the receiving means corresponds to the second keyword stored in the storage means, thereby verifying the identity of the speaker. [Prior Art Literature] [Patent Literature]

[0003] [Patent Document 1] Japanese Unexamined Patent Publication No. 2010-109618 [Summary of the Invention] [Problem to be Solved by the Invention]

[0004] Voiceprint authentication may fail to verify the speaker's identity if the audio data is short, as the authentication accuracy decreases. Therefore, Patent Document 1 performs both voiceprint authentication and speech recognition authentication to verify the speaker's identity. Consequently, the authentication device assists in verifying identity by comparing the speech recognition result obtained by speech recognition of the speaker's voice with a first keyword or a second keyword stored in the storage means, and was not intended to improve the authentication accuracy of voiceprint authentication using voiceprint information.

[0005] This disclosure was devised in view of the conventional circumstances described above, and aims to provide an authentication device and authentication method that improve the accuracy of speaker voice authentication using spoken voice. [Means for solving the problem]

[0006] This disclosure comprises: an acquisition unit that acquires an audio signal of a speaker's utterance; a detection unit that detects a first utterance section spoken by the speaker from the acquired audio signal; and an authentication unit that authenticates the speaker based on a comparison between the audio signal of the first utterance section detected by the detection unit and a database. If the authentication unit determines that the speaker cannot be authenticated, the detection unit detects a second utterance section different from the first utterance section, and the authentication unit then determines the first utterance section The audio signal and The audio signal of the second speech segment and A concatenated audio signal formed by linking these together The present invention provides an authentication device that authenticates the speaker based on a comparison with the aforementioned database.

[0007] Furthermore, this disclosure relates to an authentication method performed by one or more computers, which includes: acquiring an audio signal of a speaker's utterance; detecting a first utterance segment spoken by the speaker from the acquired audio signal; authenticating the speaker based on a comparison between the detected audio signal of the first utterance segment and a database; and, if it is determined that the speaker cannot be authenticated based on the audio signal of the first utterance segment, detecting a second utterance segment different from the first utterance segment, and the first utterance segment The audio signal and The audio signal of the second speech segment and A concatenated audio signal formed by linking these togetherThe present invention provides an authentication method for authenticating the speaker based on a comparison with the aforementioned database. [Effects of the Invention]

[0008] According to this disclosure, the accuracy of speaker voice authentication using spoken audio can be improved. [Brief explanation of the drawing]

[0009] [Figure 1] A diagram showing an example of a use case for the voice authentication system according to Embodiment 1. [Figure 2] Block diagram showing an example of the internal configuration of the recognition analysis device in Embodiment 1. [Figure 3] This diagram illustrates an example of the first user authentication process in Embodiment 1. [Figure 4] This diagram illustrates an example of a second user authentication process in Embodiment 1. [Figure 5] This diagram illustrates a third example of user authentication processing in Embodiment 1. [Figure 6] This figure illustrates a fourth example of user authentication processing in Embodiment 1. [Figure 7] A diagram illustrating a fifth user authentication process example in Embodiment 1. [Figure 8] A diagram illustrating a sixth user authentication process example in Embodiment 1. [Figure 9] A diagram illustrating a sixth user authentication process example in Embodiment 1. [Figure 10] Flowchart showing an example of the operation procedure of the recognition analysis device in Embodiment 1 [Modes for carrying out the invention]

[0010] The following describes in detail embodiments of the authentication device and authentication method disclosed herein, with reference to the drawings as appropriate. However, unnecessary details may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The accompanying drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter described in the claims.

[0011] First, with reference to Figure 1, a use case of the voice authentication system 100 according to Embodiment 1 will be described. Figure 1 is a diagram showing an example of a use case of the voice authentication system 100 according to Embodiment 1. The voice authentication system 100 acquires a voice signal or voice data from a person to be voice authenticated (in the example shown in Figure 1, user US), and compares the acquired voice signal or voice data with a number of voice signals or voice data previously registered (stored) in storage (in the example shown in Figure 1, registered speaker database DB). Based on the comparison result, the voice authentication system 100 evaluates the similarity between the user to be voice authenticated and the voice signal or voice data registered in storage, and authenticates user US based on the evaluated similarity.

[0012] The voice authentication system 100 according to Embodiment 1 comprises at least an operator-side call terminal OP1 as an example of a sound collection device, an authentication analysis device P1, a registered speaker database DB, and an information display unit DP as an example of an output device. The authentication analysis device P1 and the registered speaker database DB may be configured as a single unit. Similarly, the authentication analysis device P1 and the information display unit DP may be configured as a single unit.

[0013] Note that the voice authentication system 100 shown in FIG. 1 is an example used for authenticating a speaker (user US) in a call center, and authenticates the user US using voice data obtained by collecting the uttered voice of the user US who is talking with an operator OP. The voice authentication system 100 shown in FIG. 1 further comprises a user-side call terminal UP1 and a network NW. It goes without saying that the overall configuration of the voice authentication system 100 is not limited to the example shown in FIG. 1.

[0014] The user-side call terminal UP1 is connected to an operator-side call terminal OP1 so as to enable wireless communication via the network NW. The wireless communication referred to herein is communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark), for example.

[0015] The user-side call terminal UP1 is implemented by, for example, a notebook PC, a tablet terminal, a smartphone, a telephone, or the like. The user-side call terminal UP1 is a sound collection device provided with a microphone (not shown), which collects the uttered voice of the user US, converts the collected uttered voice into an audio signal, and transmits the converted audio signal to the operator-side call terminal OP1 via the network NW. Further, the user-side call terminal UP1 acquires the audio signal of the uttered voice of the operator OP transmitted from the operator-side call terminal OP1, and outputs the audio signal from a speaker (not shown).

[0016] The network NW is an IP network or a telephone network, and connects the user-side call terminal UP1 and the operator-side call terminal OP1 so as to enable transmission and reception of audio signals therebetween. Note that data transmission and reception are performed by wired communication or wireless communication. The wireless communication referred to herein is communication via a wireless LAN such as Wi-Fi (registered trademark), for example.

[0017] The operator-side call terminal OP1 is connected to the user-side call terminal UP1 and an authentication analysis device P1 so as to be capable of transmitting and receiving data via wired communication or wireless communication respectively, and transmits and receives audio signals.

[0018] The operator-side call terminal OP1 is implemented by, for example, a notebook PC, tablet, smartphone, or telephone. The operator-side call terminal OP1 acquires an audio signal based on the user's (US) speech transmitted from the user-side call terminal UP1 via the network NW and transmits it to the authentication analysis device P1. If the operator-side call terminal OP1 acquires an audio signal that includes both the user's (US) speech and the operator's (OP) speech, it may separate the audio signal based on the user's (US) speech from the audio signal based on the operator's (OP) speech based based on audio parameters such as the sound pressure level and frequency band of the audio signal of the operator-side call terminal OP1. After separation, the operator-side call terminal OP1 extracts only the audio signal based on the user's (US) speech and transmits it to the authentication analysis device P1.

[0019] Furthermore, the operator-side call terminal OP1 may be connected to each of the multiple user-side call terminals in a communication-enabled manner, and may simultaneously acquire voice signals from each of the multiple user-side call terminals. The operator-side call terminal OP1 transmits the acquired voice signals to the authentication analysis device P1. This allows the voice authentication system 100 to simultaneously perform voice authentication processing and voice analysis processing for each of the multiple users.

[0020] Furthermore, the operator-side call terminal OP1 may simultaneously acquire audio signals containing the individual speech voices of multiple users. The operator-side call terminal OP1 extracts individual user audio signals from the audio signals of multiple users acquired via the network NW and transmits each user's audio signal to the authentication analysis device P1. In such cases, the operator-side call terminal OP1 may analyze the audio signals of multiple users and separate and extract the audio signals for each user based on audio parameters such as sound pressure level and frequency band. If the audio signals are picked up by an array microphone or the like, the operator-side call terminal OP1 may separate and extract the audio signals for each user based on the direction of arrival of the utterance. As a result, the voice authentication system 100 can perform individual voice authentication processing and voice analysis processing for multiple users, even if the audio signals are picked up in an environment where multiple users speak simultaneously, such as a web conference.

[0021] An authentication analysis device P1, as an example of an authentication device and computer, is connected to the operator-side call terminal OP1, the registered speaker database DB, and the information display unit DP, enabling data transmission and reception. The authentication analysis device P1 may also be connected to the operator-side call terminal OP1, the registered speaker database DB, and the information display unit DP via a network (not shown) via wired or wireless communication.

[0022] The authentication analysis device P1 acquires the voice signal of user US transmitted from the operator-side call terminal OP1, and performs voice analysis on the acquired voice signal, for example, frequency by frequency, to extract the individual speech features of user US. The authentication analysis device P1 refers to the registered speaker database DB and compares the extracted speech features with the speech features of multiple users previously registered in the registered speaker database DB to perform voice authentication of user US. Alternatively, the authentication analysis device P1 may perform voice authentication of user US by comparing the extracted speech features with the speech features of a specific user previously registered in the registered speaker database DB, instead of the speech features of multiple users previously registered in the registered speaker database DB.

[0023] The authentication analysis device P1 generates an authentication result screen SC containing the user authentication result and sends it to the information display unit DP for output. It goes without saying that the authentication result screen SC shown in Figure 1 is just one example and is not limited to this. The authentication result screen SC shown in Figure 1 includes the message "Matches the voice of XXOO," which is the user authentication result.

[0024] Furthermore, the authentication analysis device P1 may perform voice authentication of user US by comparing the voice signals of multiple users pre-registered in the registered speaker database DB with the voice signal of user US. Alternatively, instead of comparing the voice signals of multiple users pre-registered in the registered speaker database DB, the authentication analysis device P1 may perform voice authentication of user US by comparing the voice signal of a specific user pre-registered in the registered speaker database DB with the voice signal of user US.

[0025] The registered speaker database DB, as an example of a database, is a so-called storage system, configured using storage media such as flash memory, HDD (Hard Disk Drive), or SSD (Solid State Drive). The registered speaker database DB stores (registers) user information and speech feature quantities of multiple users in association. User information here refers to information about the user, such as username, user ID (Identification), and identification information assigned to each user. The registered speaker database DB may be configured integrally with the authentication analysis device P1.

[0026] The information display unit DP is configured using, for example, an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display, and displays the authentication result screen SC transmitted from the authentication analysis device P1.

[0027] In the example shown in Figure 1, the user-side call terminal UP1 receives the user US's spoken voice COM12 "This is XXOO" and spoken voice COM14 "This is 123245678", converts them into audio signals, and transmits them to the operator-side call terminal OP1. The operator-side call terminal OP1 transmits the audio signals based on the user US's spoken voice COM12 and COM14 transmitted from the user-side call terminal UP1 to the authentication analysis device P1.

[0028] Furthermore, when the operator-side call terminal OP1 acquires audio signals that include the operator OP's voice COM11 "Please tell me your name" and voice COM13 "Please tell me your membership number," as well as the user US's voice COM12 and COM14, it separates and removes the audio signals based on the operator OP's voice COM11 and COM13, extracts only the audio signals based on the user US's voice COM12 and COM14, and transmits them to the authentication analysis device P1. This allows the authentication analysis device P1 to improve user authentication accuracy by using only the audio signals of the person being authenticated.

[0029] Referring to Figure 2, an example of the internal configuration of the authentication analysis device P1 will be described. Figure 2 is a block diagram showing an example of the internal configuration of the authentication analysis device P1 in Embodiment 1. The authentication analysis device P1 is composed of at least a communication unit 20, a processor 21, and a memory 22.

[0030] As an example of an acquisition unit, the communication unit 20 is connected to the operator-side call terminal OP1 and the registered speaker database DB in a manner that enables data communication. The communication unit 20 outputs the voice signal transmitted from the operator-side call terminal OP1 to the processor 21. Note that the acquisition unit is not limited to the communication unit 20, but may also be, for example, the microphone of the operator-side call terminal OP1 which is configured together with the authentication analysis device P1.

[0031] The processor 21 is composed of a semiconductor chip on which at least one of the following electronic devices is implemented: a CPU (Central Processing Unit), a DSP (Digital Signal Processor), a GPU (Graphical Processing Unit), or an FPGA (Field Programmable Gate Array). The processor 21 functions as a controller that oversees the overall operation of the authentication analysis device P1, performing control processing to coordinate the operation of each part of the authentication analysis device P1, data input / output processing between each part of the authentication analysis device P1, data calculation processing, and data storage processing.

[0032] The processor 21 implements the functions of the speech segment detection unit 21A, the speech concatenation unit 21B, the feature extraction unit 21C, and the similarity calculation unit 21D by using the programs and data stored in the ROM (Read Only Memory) 22A of the memory 22. During operation, the processor 21 uses the RAM (Random Access Memory) 22B of the memory 22 to temporarily store data or information generated or acquired by the processor 21 and each unit in the RAM 22B of the memory 22.

[0033] The speech segment detection unit 21A, as an example of a detection unit, recognition unit, conversion unit, and noise detection unit, analyzes the acquired audio signal and detects the speech segments spoken by the user US. The speech segment detection unit 21A outputs the audio signal corresponding to each speech segment detected from the audio signal (hereinafter referred to as "speech audio signal") to the speech concatenation unit 21B or the feature extraction unit 21C. The speech segment detection unit 21A may also temporarily store the speech audio signal of each speech segment in the RAM 22B of the memory 22.

[0034] As an example of a processing unit, the speech concatenation unit 21B concatenates the speech audio signals of two or more speech segments from the same person (user US) when the speech segment detection unit 21A detects these segments from the audio signal. The speech concatenation unit 21B outputs the concatenated speech audio signal (hereinafter referred to as the "concatenated audio signal") to the feature extraction unit 21C. The user authentication method will be described later.

[0035] As an example of a processing unit, the feature extraction unit 21C analyzes the characteristics of an individual's voice, for example, frequency by frequency, using one or more speech audio signals extracted by the speech interval detection unit 21A, and extracts speech features. The feature extraction unit 21C may also extract speech features from the concatenated audio signal output from the speech concatenation unit 21B. The feature extraction unit 21C outputs the extracted speech features to the similarity calculation unit 21D, or temporarily stores them in the RAM 22B of the memory 22, by associating them with the speech audio signal or concatenated audio signal from which these speech features were extracted.

[0036] As an example of an authentication unit, the similarity calculation unit 21D acquires speech features from the speech audio signal or concatenated audio signal output from the feature extraction unit 21C. The similarity calculation unit 21D refers to the registered speaker database DB and calculates the similarity between the speech features of each of the multiple users registered in the registered speaker database DB and the acquired concatenated speech features. Based on the calculated similarity, the similarity calculation unit 21D identifies the user corresponding to the speech audio signal or concatenated audio signal (i.e., the audio signal transmitted from the user-side call terminal UP1) and performs user authentication.

[0037] If the similarity calculation unit 21D determines that a user has been identified as a result of user authentication, it generates an authentication result screen SC containing information about the identified user (i.e., the authentication result) and outputs it to the information display unit DP via the display interface 23.

[0038] Furthermore, if the similarity calculation unit 21D determines that the calculated similarity is less than a predetermined value, it may determine that user authentication is not possible and generate and output a control command to the speech linking unit 21B requesting the linking of the speech audio signals. In addition, if the similarity calculation unit 21D determines that the number of times user authentication has been determined to be impossible is greater than or equal to the upper limit for user authentication for the same person (user US), it may generate an authentication result screen (not shown) notifying the user that user authentication is impossible and output it to the information display unit DP.

[0039] Memory 22 includes, for example, ROM 22A which stores a program that defines various processes performed by the processor 21 and the data used during the execution of that program, and RAM 22B which serves as work memory used when executing various processes performed by the processor 21. ROM 22A contains a program that defines various processes performed by the processor 21 and the data used during the execution of that program. RAM 22B temporarily stores data or information generated or acquired by the processor 21 (for example, pre-concatenation speech signals, concatenated speech signals, speech feature quantities corresponding to each pre-concatenation or concatenated speech segment, etc.).

[0040] The display interface 23 connects the processor 21 and the information display unit DP to enable data communication and outputs the authentication result screen SC generated by the similarity calculation unit 21D of the processor 21 to the information display unit DP.

[0041] Next, with reference to Figure 3, the first user authentication process performed by the authentication analysis device P1 will be described. Figure 3 is a diagram illustrating an example of the first user authentication process in Embodiment 1. Figures 3 to 8 show an example of a conversation between operator OP and user US, who is the target of user authentication.

[0042] The user-side call terminal UP1 records the user US's spoken words Us11 "Hello", Us12 "I don't know my PIN", Us13 "My ID is 12345678", and Us14 "My name is XXXX", converts them into audio signals, and transmits them to the operator-side call terminal OP1.

[0043] The operator-side call terminal OP1 records the operator OP's spoken words Op11 "How can I help you?", Op12 "Yes, then please tell me your ID", and Op13 "Please tell me your name", converts them into audio signals, and transmits them to the user-side call terminal UP1. The operator-side call terminal OP1 also acquires the audio signals transmitted from the user-side call terminal UP1 and transmits them to the authentication analysis device P1.

[0044] The speech segment detection unit 21A in the authentication analysis device P1 detects each speech segment of user US's speech voices Us11 to Us14 from the audio signal transmitted from the operator-side call terminal OP1. The speech segment detection unit 21A extracts the speech voice signal corresponding to each detected speech segment. In the following explanation and in Figures 3 to 8, the speech voice signal corresponding to speech voice Us11 will be referred to as "Speech 1", the speech voice signal corresponding to speech voice Us12 as "Speech 2", the speech voice signal corresponding to speech voice Us13 as "Speech 3", and the speech voice signal corresponding to speech voice Us14 as "Speech 4".

[0045] It goes without saying that the example conversations between operator OP and user US shown in Figures 3 to 8, and the voice signals used for user authentication, are examples only and are not limited thereto. The voice signals used for user authentication may be acquired by obtaining the voice signal corresponding to the spoken audio recorded after the timing of speech recognition of a predetermined word (e.g., "start") contained in the voice signal. Furthermore, the spoken audio may include multiple sentences, such as "Hello. I don't know my PIN."

[0046] The first user authentication process is described below. In the first user authentication process, if the authentication analysis device P1 determines that user authentication is not possible, it concatenates the speech audio signals corresponding to each detected speech segment in chronological order and performs user authentication again.

[0047] The feature extraction unit 21C extracts speech features from the speech audio signal "Utterance 1" corresponding to each extracted speech segment and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the speech audio signal "Utterance 1" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB and performs user authentication (first user authentication process).

[0048] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it concatenates the speech audio signal "utterance 1" and the speech audio signal "utterance 2" to the speech concatenation unit 21B. The speech concatenation unit 21B outputs the concatenated speech signal "utterance 1" + "utterance 2" to the feature extraction unit 21C. The feature extraction unit 21C extracts the speech features from the concatenated speech signal "utterance 1" + "utterance 2" and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated speech signal "utterance 1" + "utterance 2" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB and performs user authentication (second user authentication process).

[0049] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it has the speech concatenation unit 21B concatenate the speech audio signals "Speech 1", "Speech 2", and "Speech 3". The speech concatenation unit 21B outputs the concatenated speech signal "Speech 1" + "Speech 2" + "Speech 3" to the feature extraction unit 21C. The feature extraction unit 21C extracts the speech features from the concatenated speech signals "Speech 1" + "Speech 2" + "Speech 3" and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated speech signals "Speech 1" + "Speech 2" + "Speech 3" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB, and performs user authentication (third user authentication process).

[0050] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it has the speech concatenation unit 21B concatenate the speech audio signals "Speech 1", "Speech 2", "Speech 3", and "Speech 4". The speech concatenation unit 21B outputs the concatenated speech signals "Speech 1" + "Speech 2" + "Speech 3" + "Speech 4" to the feature extraction unit 21C. The feature extraction unit 21C extracts speech features from the concatenated speech signals "Speech 1" + "Speech 2" + "Speech 3" + "Speech 4" and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated speech signal "utterance 1" + "utterance 2" + "utterance 3" + "utterance 4" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB, and performs user authentication (4th user authentication process).

[0051] As described above, the authentication analysis device P1 performs user authentication using the speech audio signal corresponding to each utterance. If it determines that user authentication is not possible, it sequentially concatenates the speech audio signals in chronological order, increasing the signal length (utterance length) of the concatenated audio signal used in the user authentication process. This allows the individuality of the speech features of each user US to be more strongly expressed.

[0052] As a result, the authentication analysis device P1 according to Embodiment 1 can improve user authentication accuracy because, even if there is variation in the speech features of the user US included in each speech audio signal, the individuality of the speech features used for user authentication is more strongly expressed.

[0053] Furthermore, this allows the authentication analysis device P1 according to Embodiment 1 to repeatedly perform user authentication using the speech audio signals of each utterance segment detected from the acquired audio signal. Therefore, if user US is authenticated during a call (conversation) between operator OP and user US, operator OP can end the call (conversation) with user US more quickly.

[0054] In the example shown in Figure 3, the user authentication process is described as being performed four times. However, the authentication analysis device P1 may terminate the user authentication process when it determines that user authentication has been successful. Furthermore, the authentication analysis device P1 may have an upper limit set on the number of times the user authentication process has been performed. If it determines that the upper limit has been reached, it may generate an authentication result screen (not shown) indicating that user authentication is not possible and output it to the information display unit DP.

[0055] Next, with reference to Figure 4, the second user authentication process performed by the authentication analysis device P1 will be described. Figure 4 is a diagram illustrating an example of the second user authentication process in Embodiment 1.

[0056] In the second user authentication process, the authentication analysis device P1 concatenates multiple speech voice signals used for user authentication so that the signal length is equal to or greater than a predetermined time (e.g., 5 seconds, 10 seconds, etc.), and performs user authentication using the concatenated voice signal. In the example shown in Figure 4, an example where the predetermined time is 10 seconds is described, but it goes without saying that the predetermined time is not limited to this.

[0057] In the example shown in Figure 4, the speech segment detection unit 21A detects each of the speech audio signals "Speech 1" to "Speech 4" corresponding to each speech segment and outputs them to the speech linking unit 21B. In Figure 4, the signal length of speech audio signal "Speech 1" is 0.8 seconds, the signal length of speech audio signal "Speech 2" is 2.9 seconds, the signal length of speech audio signal "Speech 3" is 4.0 seconds, and the signal length of speech audio signal "Speech 4" is 3.5 seconds.

[0058] The speech concatenation unit 21B combines and concatenates the speech audio signals "Speech 1" to "Speech 4" so that the total signal length of the speech audio signals used for user authentication is equal to or greater than a predetermined time. If the signal length of a single speech audio signal is equal to or greater than a predetermined time, the speech concatenation process by the speech concatenation unit 21B may be omitted. The speech concatenation unit 21B outputs the concatenated audio signal to the feature extraction unit 21C.

[0059] The feature extraction unit 21C acquires a speech audio signal or concatenated speech signal having a signal length of a predetermined time or longer, output from the speech interval detection unit 21A or the speech concatenation unit 21B. The feature extraction unit 21C extracts the speech features of user US included in the acquired speech audio signal or concatenated speech signal. The feature extraction unit 21C outputs the extracted speech features of user US to the similarity calculation unit 21D.

[0060] The similarity calculation unit 21D acquires speech features from the speech audio signal or concatenated speech signal output from the feature extraction unit 21C. The similarity calculation unit 21D refers to the registered speaker database DB and calculates the similarity between the acquired speech features and the speech features of each of the multiple users registered in the registered speaker database DB. Based on the calculated similarity, the similarity calculation unit 21D identifies the user corresponding to the acquired speech audio signal or concatenated speech signal and performs user authentication.

[0061] For example, in the example shown in Figure 4, the signal length of the concatenated audio signal "Utterance 1" + "Utterance 2", which is formed by concatenating the speech audio signals "Utterance 1" and "Utterance 2", is 3.7 seconds (i.e., less than the predetermined time (10 seconds)). In the second user authentication process, user authentication is not performed using speech audio signals whose concatenated signal length is less than the predetermined time.

[0062] Furthermore, the combined voice signal "Utterance 1" + "Utterance 2" + "Utterance 3" + "Utterance 4", formed by concatenating the speech voice signals "Utterance 1" to "Utterance 4", has a signal length of 11.2 seconds (i.e., at least the predetermined time (10 seconds)). Similarly, the combined voice signal "Utterance 3" + "Utterance 4" + "Utterance 2", formed by concatenating the speech voice signals "Utterance 2" to "Utterance 4", has a signal length of 10.4 seconds (i.e., at least the predetermined time (10 seconds)). In such cases, the authentication analysis device P1 performs user authentication processing using the combined voice signal "Utterance 1" + "Utterance 2" + "Utterance 3" + "Utterance 4", or the combined voice signal "Utterance 3" + "Utterance 4" + "Utterance 2".

[0063] If the authentication analysis device P1 determines that user authentication is not possible, it generates a new concatenated speech signal using a different combination of speech signals than the one already used for user authentication, and attempts user authentication again. For example, if the authentication analysis device P1 performs the first user authentication process using the concatenated speech signal "utterance 3" + "utterance 4" + "utterance 2" and determines that user authentication is not possible, it performs the second user authentication process using the concatenated speech signal "utterance 1" + "utterance 2" + "utterance 3" + "utterance 4".

[0064] In the second user authentication process, the concatenation order of the spoken audio signals may be in chronological order, such as "Speech 1" + "Speech 2" + "Speech 3" + "Speech 4", or it may be in order of the length of the spoken audio signals, such as "Speech 3" + "Speech 4" + "Speech 2".

[0065] Furthermore, in the second user authentication process, the speech concatenation unit 21B may select the speech audio signals to be concatenated. If a lower time limit (for example, 2 seconds) is set as a criterion for selecting the speech audio signals to be concatenated, the speech concatenation unit 21B may determine whether the signal length of the speech audio signal corresponding to each speech segment output from the speech segment detection unit 21A is equal to or greater than the lower time limit. The speech concatenation unit 21B then performs the speech audio signal concatenation process using the speech audio signals whose signal length is determined to be equal to or greater than the lower time limit.

[0066] As a result, the authentication analysis device P1 can remove short utterances such as "yes" or "uh-huh," which have small individual user US speech features, from the speech signal used for user authentication. Therefore, the authentication analysis device P1 can perform user authentication using concatenated speech signals that contain speech features that more strongly reflect individuality, thereby improving user authentication accuracy.

[0067] As described above, the authentication analysis device P1 in Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features included in each speech signal, by using concatenated speech signals that have a signal length of a predetermined time or longer and have speech features more suitable for user authentication processing.

[0068] Next, with reference to Figure 5, the third user authentication process performed by the authentication analysis device P1 will be described. Figure 5 is a diagram illustrating an example of the third user authentication process in Embodiment 1.

[0069] In the third user authentication process, the authentication analysis device P1 recognizes the number of characters contained in the speech signal used for user authentication, concatenates multiple speech signals so that the recognized number of characters is equal to or greater than a predetermined number of characters (e.g., 20 characters, 25 characters, etc.), and performs user authentication using the concatenated speech signal. In the example shown in Figure 5, an example where the predetermined number of characters is 25 characters is described, but it goes without saying that the predetermined time is not limited to this. Note that the number of characters here may be the number of morae, syllables, phonemes, etc.

[0070] In the example shown in Figure 5, the speech segment detection unit 21A detects each of the speech audio signals "Speech 1" to "Speech 4" corresponding to each speech segment, recognizes the number of characters contained in each speech audio signal, and outputs the recognition result and the speech audio signal to the speech concatenation unit 21B. In Figure 5, the speech audio signal "Speech 1" has 5 characters, the speech audio signals "Speech 2" and "Speech 3" each have 16 characters, and the speech audio signal "Speech 4" has 12 characters.

[0071] The speech concatenation unit 21B combines and concatenates the speech audio signals "Speech 1" to "Speech 4" so that the total number of characters in the speech audio signals used for user authentication is equal to or greater than a predetermined number of characters. If the number of characters in a single speech audio signal is equal to or greater than the predetermined number of characters, the speech concatenation process by the speech concatenation unit 21B may be omitted. The speech concatenation unit 21B outputs the concatenated audio signal to the feature extraction unit 21C.

[0072] The feature extraction unit 21C acquires a speech audio signal or concatenated speech audio signal containing a predetermined number of characters or more, output from the speech segment detection unit 21A or the speech concatenation unit 21B. The feature extraction unit 21C extracts the speech features of user US included in the acquired speech audio signal or concatenated speech audio signal. The feature extraction unit 21C outputs the extracted speech features of user US to the similarity calculation unit 21D.

[0073] The similarity calculation unit 21D acquires speech features from the speech audio signal or concatenated speech signal output from the feature extraction unit 21C. The similarity calculation unit 21D refers to the registered speaker database DB and calculates the similarity between the speech features of each of the multiple users registered in the registered speaker database DB and the acquired concatenated speech features. Based on the calculated similarity, the similarity calculation unit 21D performs user authentication.

[0074] For example, in the example shown in Figure 5, the number of characters in the concatenated speech signal "Utterance 1" + "Utterance 2", which is formed by concatenating the speech signal "Utterance 1" and the speech signal "Utterance 2", is 21 characters (i.e., less than the predetermined number of characters (25 characters)). In the third user authentication process, user authentication is not performed using a concatenated speech signal in which the number of characters after concatenation is less than the predetermined number of characters.

[0075] Furthermore, the concatenated audio signal "Utterance 1" + "Utterance 2" + "Utterance 3" + "Utterance 4", formed by concatenating the speech signals "Utterance 1" to "Utterance 4", has 49 characters (i.e., at least the predetermined number of characters (25 characters)). Similarly, the concatenated audio signal "Utterance 3" + "Utterance 4" + "Utterance 2", formed by concatenating the speech signals "Utterance 2" to "Utterance 4", has 44 characters (i.e., at least the predetermined number of characters (25 characters)). The authentication analysis device P1 performs user authentication processing using the concatenated audio signal "Utterance 1" + "Utterance 2" + "Utterance 3" + "Utterance 4", or the concatenated audio signal "Utterance 3" + "Utterance 4" + "Utterance 2".

[0076] If the authentication analysis device P1 determines that user authentication is not possible, it will attempt user authentication again using a new concatenated voice signal with a different combination of speech signals than those already used for user authentication. For example, if the authentication analysis device P1 performs the first user authentication process using the concatenated voice signal "utterance 3" + "utterance 4" + "utterance 2" and determines that user authentication is not possible, it will perform the second user authentication process using the concatenated voice signal "utterance 1" + "utterance 2" + "utterance 3" + "utterance 4".

[0077] In the third user authentication process, the concatenation order of the spoken audio signals may be chronological, such as "Speech 1" + "Speech 2" + "Speech 3" + "Speech 4", or it may be in descending order of the number of characters in the spoken audio signals, such as "Speech 3" + "Speech 4" + "Speech 2".

[0078] Furthermore, in the third user authentication process, the speech concatenation unit 21B may select speech audio signals to be concatenated. If a minimum number of characters (for example, 5 characters) is set as a criterion for selecting speech audio signals to be concatenated, the speech concatenation unit 21B may determine whether the number of characters in the speech audio signal corresponding to each speech segment output from the speech segment detection unit 21A is equal to or greater than the minimum number of characters. The speech concatenation unit 21B then performs the speech audio signal concatenation process using the speech audio signal whose signal length is determined to be equal to or greater than the minimum number of characters.

[0079] As a result, the authentication analysis device P1 can remove speech signals with a small number of characters, such as "yes" or "uh-huh," which have small individual speech features, from the speech signals used for user authentication. Therefore, the authentication analysis device P1 can perform user authentication using speech signals or concatenated speech signals that contain speech features that more strongly express individuality, thereby improving user authentication accuracy.

[0080] As described above, the authentication analysis device P1 in Embodiment 1 can perform user authentication processing using a speech audio signal or concatenated speech signal that contains a number of characters greater than or equal to a predetermined number and has speech features suitable for user authentication processing.

[0081] As a result, the authentication analysis device P1 according to Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features included in each speech signal.

[0082] Next, with reference to Figure 6, the fourth user authentication process performed by the authentication analysis device P1 will be described. Figure 6 is a diagram illustrating an example of the fourth user authentication process in Embodiment 1.

[0083] In the fourth user authentication process, the authentication analysis device P1 performs weighting on each speech signal based on the number of characters in the speech signal. The authentication analysis device P1 then performs user authentication using the speech features after weighting.

[0084] In the example shown in Figure 6, the speech segment detection unit 21A detects each of the speech audio signals "Speech 1" to "Speech 4" corresponding to each speech segment, performs speech recognition to determine the number of characters contained in each speech audio signal, and outputs the speech recognition result and the speech audio signal to the speech concatenation unit 21B. In Figure 6, the speech audio signal "Speech 1" has 5 characters, the speech audio signals "Speech 2" and "Speech 3" each have 16 characters, and the speech audio signal "Speech 4" has 12 characters.

[0085] The speech concatenation unit 21B determines weight coefficients for each speech signal based on the speech signal recognized by the speech segment detection unit 21A and the number of characters in each speech signal. The speech concatenation unit 21B concatenates the speech signals to generate a concatenated speech signal and outputs it to the feature extraction unit 21C.

[0086] Specifically, the speech concatenation unit 21B calculates the total number of characters in two or more speech signals to be concatenated, calculates the ratio of the number of characters in each speech signal to the calculated total number of characters, and determines a weighting coefficient corresponding to the calculated ratio. The weighting coefficients corresponding to each speech segment may also be output to and stored in the RAM 22B.

[0087] The feature extraction unit 21C performs weighting on the speech features extracted from each utterance segment based on each of the two or more speech segments contained in the concatenated speech signal output from the speech concatenation unit 21B, and the weight coefficients corresponding to each speech segment. If the user authentication process is performed for the first time and a concatenated speech signal is not generated, the calculation of the weight coefficients and the weighting process may be performed by the speech segment detection unit 21A, or the process itself may be omitted.

[0088] The following is a specific example of the fourth user authentication process, with reference to Figure 6.

[0089] The speech concatenation unit 21B determines a weight coefficient of 1.0 based on the number of characters (5 characters) in the speech-recognized speech signal "Utterance 1" and the total number of characters in the speech signal used for the first user authentication process (i.e., the speech signal "Utterance 1"). The speech concatenation unit 21B outputs the speech signal and the weight coefficient to the feature extraction unit 21C.

[0090] The feature extraction unit 21C extracts speech features from the speech audio signal "utterance 1" output from the speech concatenation unit 21B, weights the extracted speech features of the speech audio signal "utterance 1" with weight coefficients, and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the speech audio signal "utterance 1" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB, and performs user authentication (first user authentication process).

[0091] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it instructs the speech concatenation unit 21B to concatenate the speech audio signal "Speech 1" and the speech audio signal "Speech 2". The speech concatenation unit 21B determines the respective weight coefficients for speech audio signals "Speech 1" and "Speech 2" based on the number of characters in speech audio signal "Speech 1" (5 characters) and the number of characters in speech audio signal "Speech 2" (16 characters), and the sum of the number of characters in these speech audio signals (5 + 16). In the example shown in Figure 6, the speech concatenation unit 21B determines the weight coefficient for speech audio signal "Speech 1" to be 0.24 and the weight coefficient for speech audio signal "Speech 2" to be 0.76. The speech concatenation unit 21B outputs the concatenated audio signal and each weight coefficient to the feature extraction unit 21C.

[0092] The feature extraction unit 21C extracts speech features from the speech audio signals "Utterance 1" and "Utterance 2" output from the speech concatenation unit 21B. The feature extraction unit 21C weights the extracted speech features of each speech audio signal "Utterance 1" and "Utterance 2" with corresponding weight coefficients and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated audio signals "Utterance 1" + "Utterance 2" output from the feature extraction unit 21C with the speech features of multiple users registered in the registered speaker database DB and performs user authentication (second user authentication process).

[0093] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it has the speech concatenation unit 21B concatenate the speech audio signals "Speech 1", "Speech 2", and "Speech 3". The speech concatenation unit 21B determines the weight coefficients for each of the speech audio signals "Speech 1", "Speech 2", and "Speech 3" based on the number of characters in "Speech 1" (5 characters), the number of characters in "Speech 2" (16 characters), and the number of characters in "Speech 3" (16 characters), and the sum of the number of characters in these speech audio signals (5 + 16 + 16). In the example shown in Figure 6, the speech concatenation unit 21B determines the weight coefficient for "Speech 1" to be 0.14, and the weight coefficients for "Speech 2" and "Speech 3" to be 0.43, respectively. The speech concatenation unit 21B outputs the concatenated speech signal and each weight coefficient to the feature extraction unit 21C.

[0094] The feature extraction unit 21C extracts speech features from the speech audio signals "Speech 1", "Speech 2", and "Speech 3" output from the speech concatenation unit 21B. The feature extraction unit 21C weights each of the extracted speech features of speech audio signals "Speech 1" to "Speech 3" by assigning corresponding weight coefficients and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated audio signals "Speech 1" + "Speech 2" + "Speech 3" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB and performs user authentication (third user authentication process).

[0095] If the similarity calculation unit 21D determines that user authentication is not possible based on the calculated similarity, it has the speech concatenation unit 21B concatenate the speech audio signals "Speech 1", "Speech 2", "Speech 3", and "Speech 4". The speech concatenation unit 21B determines the respective weight coefficients for speech audio signals "Speech 1" and "Speech 2" based on the number of characters in speech audio signals "Speech 1" (5 characters), "Speech 2" (16 characters), "Speech 3" (16 characters), and "Speech 4" (12 characters), and the sum of the number of characters in these speech audio signals (5 + 16 + 16 + 12). In the example shown in Figure 6, the speech concatenation unit 21B determines the weight coefficient for the speech audio signal "Speech 1" to be 0.10, the weight coefficients for the speech audio signals "Speech 2" and "Speech 3" to be 0.33, and the weight coefficient for the speech audio signal "Speech 4" to be 0.24. The speech concatenation unit 21B outputs the concatenated audio signals and each weight coefficient to the feature extraction unit 21C.

[0096] The feature extraction unit 21C extracts speech features from the speech audio signals "utterance 1", "utterance 2", "utterance 3", and "utterance 4" output from the speech concatenation unit 21B. The feature extraction unit 21C weights each of the extracted speech features of speech audio signals "utterance 1" to "utterance 4" by assigning corresponding weight coefficients and outputs them to the similarity calculation unit 21D. The similarity calculation unit 21D compares the speech features of the concatenated audio signals "utterance 1" + "utterance 2" + "utterance 3" + "utterance 4" output from the feature extraction unit 21C with the speech features of each of the multiple users registered in the registered speaker database DB and performs user authentication (4th user authentication process).

[0097] In the fourth user authentication process example described above, an example of determining weighting coefficients based on the number of characters was explained, but the method is not limited to this. For example, the weighting coefficients may be determined based on the number of morae, syllables, or phonemes. Furthermore, it goes without saying that the above example of calculating weighting coefficients is just one example and the method is not limited to this.

[0098] As described above, the authentication analysis device P1 in Embodiment 1 can perform user authentication processing using a speech audio signal that has speech features more suitable for user authentication processing by weighting the speech features of the speech audio signal.

[0099] As a result, the authentication analysis device P1 according to Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features included in each speech signal.

[0100] Next, with reference to Figure 7, a fifth user authentication process performed by the authentication analysis device P1 will be described. Figure 7 is a diagram illustrating an example of the fifth user authentication process in Embodiment 1.

[0101] In the fifth user authentication process, the speech segment detection unit 21A of the authentication analysis device P1 analyzes the speech audio signal and detects segments containing noise (e.g., voices other than the user US, background noise, ambient noise, etc.) in the speech audio signal (hereinafter referred to as "noise segments"). The speech segment detection unit 21A either deletes the detected noise segments from the speech audio signal or deletes the speech audio signal itself corresponding to the speech segment containing the noise segment from the concatenated audio signal. The authentication analysis device P1 then performs the user authentication process using the speech audio signal or concatenated audio signal after the deletion process.

[0102] The spoken audio Us12 shown in Figure 7 includes noise Nz11 "ding-dong," which is the ambient sound of the user US. In such cases, the speech section detection unit 21A detects each of the spoken audio signals "Speech 1" to "Speech 4" corresponding to each spoken section, detects the noise Nz11 from the concatenated audio signal obtained by concatenating each of the detected spoken audio signals "Speech 1" to "Speech 4," and detects the noise section Nz that contains this noise Nz11.

[0103] The speech segment detection unit 21A removes the noise segment Nz detected from the speech audio signal "Speech 2", and generates a concatenated audio signal by concatenating the speech audio signal "Speech 2" after removing the noise segment Nz with the speech audio signals "Speech 1", "Speech 3", and "Speech 4" corresponding to each speech segment.

[0104] Furthermore, the speech segment detection unit 21A deletes the speech audio signal "Speech 2" which contains the noise segment Nz, and generates a concatenated audio signal by concatenating the speech audio signals "Speech 1", "Speech 3", and "Speech 4", which do not contain the noise segment Nz.

[0105] Here, we will describe an example in which the speech segment detection unit 21A detects and removes the noise segment Nz from the concatenated speech signal, but the same procedure applies when detecting and removing the noise segment Nz from the speech signal.

[0106] As described above, the authentication analysis device P1 in Embodiment 1 can perform user authentication processing using a speech audio signal that has speech features more suitable for user authentication processing by removing noise contained in the speech audio signal. As a result, the authentication analysis device P1 according to Embodiment 1 can improve the accuracy of user authentication.

[0107] Next, the sixth user authentication process performed by the authentication analysis device P1 will be described with reference to Figures 8 and 9. Figure 8 is a diagram illustrating an example of the sixth user authentication method in Embodiment 1. Figure 9 is a diagram illustrating an example of the sixth user authentication method in Embodiment 1.

[0108] In the sixth user authentication process, the speech segment detection unit 21A of the authentication analysis device P1 analyzes the spoken audio signal to recognize the number of characters and calculates the speech rate of this spoken audio signal (i.e., the number of characters per second). The speech segment detection unit 21A performs a process to reduce or extend the spoken audio signal so that the speech rate of the spoken audio signal becomes a predetermined speech rate (hereinafter referred to as "speech rate conversion process"). For example, in the example shown in Figure 9, the spoken audio signal Dt1 is converted to the spoken audio signal Dt2 by the speech rate conversion process. The authentication analysis device P1 performs user authentication using the spoken audio signal after the speech rate conversion process, or a concatenated audio signal obtained by concatenating the spoken audio signals after the speech rate conversion process.

[0109] If the speech speeds of the source data (i.e., the speech audio signals) from which the speech features of multiple users registered (stored) in the registered speaker database DB are the same (for example, the speech speed shown in Figure 8 = 5.0 characters / second), the speech section detection unit 21A sets this same speech speed as the predetermined speech speed and executes the speech speed conversion process. As a result, the authentication analysis device P1 can calculate the similarity between the speech features of the speech audio signal or concatenated audio signal used for user authentication and the speech features of each user registered in the registered speaker database DB with higher accuracy, thereby improving the accuracy of user authentication.

[0110] The following section will specifically explain examples of speech rate conversion processing for each of the user US's speech signals, "Speech 1" to "Speech 4," with reference to Figure 8.

[0111] For example, when registering user US's voice (speech features), the speech signal of user US used for registration (storage) in the registered speaker database DB has 17 characters, a speech duration (i.e., speech interval) of 3.6 seconds, the content of the speech is "Please register my voice," and the speech rate is 4.72 characters / second. In this case, the speech signal of user US with a speech rate of 4.72 characters / second is registered (storage) in the registered speaker database DB after undergoing a speech rate conversion process that expands it to a speech signal with a predetermined speech rate of 5.0 characters / second. Note that the speech rate conversion process during registration (storage) in the registered speaker database DB may be performed by the authentication analysis device P1.

[0112] During user authentication, user US's spoken voice signal "Speech 1" consists of 5 characters, is spoken for 0.8 seconds, says "Hello", and is spoken at a speed of 6.25 characters / second. Speech voice signal "Speech 2" consists of 16 characters, is spoken for 2.9 seconds, says "I don't know my PIN", and is spoken at a speed of 5.51 characters / second. Speech voice signal "Speech 3" consists of 16 characters, is spoken for 4.0 seconds, says "My ID is 12345678", and is spoken at a speed of 4.0 characters / second. Speech voice signal "Speech 4" consists of 12 characters, is spoken for 3.5 seconds, says "My name is XXOO", and is spoken at a speed of 3.42 characters / second.

[0113] Each of the speech signals "Utterance 1" to "Utterance 4" is converted to a predetermined speech speed of 5.0 characters / second when registered (stored) in the registered speaker database DB. As a result, speech signal "Utterance 1" is converted to a speech signal with a speech duration of 1.0 second. Similarly, speech signals "Utterance 2" and "Utterance 3" are converted to speech signals with a speech duration of 3.2 seconds each. Speech signal "Utterance 4" is converted to a speech signal with a speech duration of 2.4 seconds.

[0114] The speech rate of the spoken audio signal may be calculated based on the number of characters and the duration of speech obtained from the speech recognition results of the spoken audio signal, or it may be estimated based on the number of morae, syllables, or phonemes and the duration of speech. Alternatively, the speech rate of the spoken audio signal may be estimated directly from the temporal and frequency components of the audio signal through computational processing.

[0115] As described above, even when there is variation in the speech rate of the user US, the authentication analysis device P1 in Embodiment 1 can perform user authentication processing using the speech audio signal converted to a predetermined speech rate. This allows for more accurate calculation of the similarity between the speech feature quantities of the speech audio signal or concatenated speech signal used for user authentication and the speech feature quantities for each user registered in the registered speaker database DB, thereby improving user authentication accuracy.

[0116] Next, with reference to Figure 10, an example of the operation procedure of the authentication analysis device P1 will be described. Figure 10 is a flowchart showing an example of the operation procedure of the authentication analysis device P1 in Embodiment 1.

[0117] The communication unit 20 in the authentication analysis device P1 acquires the voice signal (or voice data) transmitted from the operator-side call terminal OP1 (St11). The communication unit 20 outputs the acquired voice signal to the processor 21.

[0118] When the processor 21 acquires the audio signal output from the communication unit 20, it starts authenticating the user US, who is the target of the audio authentication of the acquired audio signal (St12).

[0119] The speech segment detection unit 21A in the processor 21 detects a speech segment from the acquired audio signal (St13).

[0120] The speech segment detection unit 21A recognizes the number of characters contained in the speech audio signal corresponding to the speech segment. Based on the recognized number of characters and the signal length of the speech audio signal (speech length, speech duration, etc.), the speech segment detection unit 21A calculates the speech rate of the speech audio signal. The speech segment detection unit 21A performs a speech rate conversion process on the speech audio signal and converts the speech rate of the speech audio signal to a predetermined speech rate (St14). Note that the process in step St14 is not mandatory and may be omitted.

[0121] The speech segment detection unit 21A stores information about the detected speech segment (for example, the start and end times of the speech segment, the number of characters, the signal length (speech length, speech duration, etc.), the speech speed before or after speech speed conversion, etc.) in the memory 22 (St15).

[0122] The speech segment detection unit 21A selects one or more speech audio signals to be used for user authentication based on the currently set user authentication processing method (St16). Although not shown in Figure 10, if the authentication analysis device P1 determines that there are no speech audio signals to be used for user authentication based on the currently set user authentication processing method, it may return to the process in step St13 to detect a new speech segment.

[0123] The speech segment detection unit 21A performs a speech concatenation process to concatenate each of the selected speech audio signals and generates a concatenated audio signal (St17). Note that the process in step St17 is omitted if the first user authentication method is set and before the first user authentication is performed. The speech segment detection unit 21A outputs the generated concatenated audio signal to the feature extraction unit 21C.

[0124] The feature extraction unit 21C extracts individual user US speech features from the concatenated speech signal output from the speech segment detection unit 21A (St18). The feature extraction unit 21C outputs the extracted individual user US speech features to the similarity calculation unit 21D.

[0125] The similarity calculation unit 21D refers to the speech feature quantities of each of the multiple users registered in the registered speaker database DB and calculates the similarity between the speech feature quantities of individual user US output from the feature extraction unit 21C and the speech feature quantities of each of the multiple users registered in the registered speaker database DB (St19).

[0126] The similarity calculation unit 21D determines whether there are any users among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than a threshold (St20).

[0127] In step St19, if the similarity calculation unit 21D determines that there is a user among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than a threshold (St20, YES), it determines that this user is user US of the speech signal (St21). If the similarity calculation unit 21D determines that there are multiple users whose similarity is equal to or greater than a threshold, it may determine that the user with the highest similarity is user US of the speech signal.

[0128] If the similarity calculation unit 21D determines that a user has been identified, it generates an authentication result screen SC containing information about the identified user (i.e., the authentication result) and outputs it to the information display unit DP via the display I / F 23 (St23).

[0129] On the other hand, in the process of step St19, if the similarity calculation unit 21D determines that there are no users among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than the threshold (St20, NO), it determines whether the current number of user authentication processes is equal to or greater than the set upper limit (St22).

[0130] In step St22, if the similarity calculation unit 21D determines that the current number of user authentication attempts is greater than or equal to the set upper limit (St22, YES), it determines, based on the acquired audio signal, that user authentication is not possible (i.e., user authentication has failed) (St24). The similarity calculation unit 21D generates an authentication result screen (not shown) notifying that user authentication is not possible and transmits it to the information display unit DP via the display I / F 23. The information display unit DP outputs (displays) the authentication result screen transmitted from the authentication analysis device P1.

[0131] If the similarity calculation unit 21D determines in step St22 that the current number of user authentication processes is not equal to or greater than the set upper limit (St22, NO), it returns to step St13.

[0132] As described above, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using a speech voice signal more suitable for user authentication processing by a predetermined user authentication processing method. This makes it possible to improve the user authentication accuracy of the authentication analysis device P1 according to Embodiment 1.

[0133] As described above, the authentication analysis device P1 according to Embodiment 1 comprises: a communication unit 20 (an example of an acquisition unit) that acquires the voice signal of a speaker's (e.g., user US, etc.) speech; a speech segment detection unit 21A (an example of a detection unit) that detects a first speech segment spoken by the speaker from the acquired voice signal; and a similarity calculation unit 21D (an example of an authentication unit) that authenticates the speaker (i.e., authenticates the user) based on a comparison between the speech voice signal (an example of a voice signal) of the first speech segment detected by the speech segment detection unit 21A and a registered speaker database DB (an example of a database). If the similarity calculation unit 21D determines that the speaker cannot be authenticated, the speech segment detection unit 21A detects a second speech segment different from the first speech segment. The similarity calculation unit 21D authenticates the speaker based on a comparison between the speech voice signals of the first and second speech segments and the registered speaker database DB. Note that one or more computers are configured to include at least the authentication analysis device P1.

[0134] As a result, if the authentication analysis device P1 according to Embodiment 1 determines that user authentication cannot be performed using the speech audio signal of one utterance section (first utterance section), it sequentially concatenates the speech audio signals in chronological order, and by increasing the signal length (utterance length) of the concatenated audio signal used for user authentication processing, it can extract speech features that more strongly express individuality. Therefore, even if there is variation in the user's speech features included in each speech audio signal, the authentication analysis device P1 according to Embodiment 1 can extract speech features that more strongly express individuality for use in user authentication, thereby improving user authentication accuracy.

[0135] Furthermore, the speech segment detection unit 21A in the authentication analysis device P1 according to Embodiment 1 detects the first speech segment and the second speech segment respectively in accordance with the time series of the acquired audio signal. As a result, the authentication analysis device P1 according to Embodiment 1 can re-execute the user authentication process using the speech audio signals of the multiple speech segments that were sequentially detected in accordance with the time series of the audio signal.

[0136] Furthermore, in Embodiment 1, the first and second speech segments are two consecutive speech segments detected by the speech segment detection unit 21A. This allows the system to determine that user authentication is not possible using the speech audio signal of a single speech segment (i.e., the first speech segment). By sequentially concatenating the speech audio signals in chronological order and increasing the signal length (speech length) of the concatenated audio signal used for user authentication, the system can extract speech features that more strongly reflect individual characteristics. As a result, the authentication analysis device P1 according to Embodiment 1 can extract speech features that more strongly reflect individual characteristics for use in user authentication, even if there is variation in the user's speech features contained in each speech audio signal, thereby improving user authentication accuracy.

[0137] Furthermore, in Embodiment 1, the total length of the first utterance segment and the second utterance segment is equal to or greater than a first predetermined time (for example, 5 seconds or more). As a result, the authentication analysis device P1 according to Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features included in each utterance speech signal by using concatenated speech signals with a signal length equal to or greater than the first predetermined time.

[0138] Furthermore, in the authentication analysis device P1 according to Embodiment 1, the length of each of the first and second utterance segments is equal to or greater than a second predetermined time (for example, 10 seconds or more). As a result, the authentication analysis device P1 according to Embodiment 1 can remove speech signals that contain short utterances such as "yes" or "uh-huh," which have small individual speech features, from the speech signals used for user authentication. Therefore, the authentication analysis device P1 can perform user authentication using concatenated speech signals that contain speech features that more strongly express individuality, thereby improving user authentication accuracy. In addition, the authentication analysis device P1 in Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features contained in each speech signal, by using concatenated speech signals that have a signal length of a predetermined time or more and contain speech features more suitable for user authentication processing.

[0139] Furthermore, the authentication analysis device P1 according to Embodiment 1 further includes a speech segment detection unit 21A (an example of a recognition unit) that performs speech recognition of a first number of characters included in a first speech segment and a second number of characters included in a second speech segment. The total number of characters included in the first and second speech segments is equal to or greater than a first predetermined number of characters (for example, 25 characters). As a result, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using a speech audio signal or concatenated speech signal that contains a number of characters equal to or greater than the predetermined number and has speech features more suitable for user authentication processing. Therefore, the authentication analysis device P1 can perform user authentication using a speech audio signal or concatenated speech signal that includes speech features that more strongly express individuality, thereby improving user authentication accuracy.

[0140] Furthermore, in the authentication analysis device P1 according to Embodiment 1, the number of characters included in the first utterance section and the second utterance section is at least a second predetermined number of characters (for example, 5 characters). As a result, the authentication analysis device P1 according to Embodiment 1 can remove utterances with a small number of characters, such as "yes" and "uh-huh," which have small utterance features for individual user US, from the utterance audio signal used for user authentication. Therefore, the authentication analysis device P1 can perform user authentication using an utterance audio signal or concatenated audio signal that includes utterance features that more strongly express individuality, thereby improving user authentication accuracy.

[0141] Furthermore, the authentication analysis device P1 according to Embodiment 1 further includes a speech segment detection unit 21A that performs speech recognition to determine the number of characters in a first speech segment and the number of characters in a second speech segment. The similarity calculation unit 21D performs weighting on the speech audio signal of the first speech segment based on the number of characters in the first speech segment and on the speech audio signal of the second speech segment based on the number of characters in the second speech segment, and authenticates the speaker based on a comparison between the weighted speech audio signals of the first and second speech segments and the registered speaker database DB. As a result, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using speech audio signals with speech features more suitable for user authentication processing by performing weighting processing on each speech audio signal based on the ratio of the number of characters in each speech audio signal to the total number of characters in the concatenated speech signals used for user authentication processing. Therefore, the authentication analysis device P1 according to Embodiment 1 can improve user authentication accuracy even if there is variation in the user's speech features included in each speech audio signal.

[0142] Furthermore, the authentication analysis device P1 according to Embodiment 1 further includes a speech concatenation unit 21B and a feature extraction unit 21C (an example of a processing unit) that weight the first and second speech segments based on the number of first and second characters recognized by the speech segment detection unit 21A. The speech concatenation unit 21B calculates the total number of characters based on the number of first and second characters, and performs weighting on the first speech segment based on the ratio of the number of first characters to the total number of characters, and weighting on the second speech segment based on the ratio of the number of second characters to the total number of characters. The similarity calculation unit 21D authenticates the speaker based on a comparison between the speech signals of the weighted first and second speech segments and the registered speaker database DB. As a result, even when there is variation in the speech rate of the user US, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using the speech audio signal converted to a predetermined speech rate. This allows for more accurate calculation of the similarity between the speech feature quantities of the speech audio signal or concatenated speech signal used for user authentication and the speech feature quantities for each user registered in the registered speaker database DB, thereby improving the accuracy of user authentication.

[0143] Furthermore, the authentication analysis device P1 according to Embodiment 1 further includes a speech segment detection unit 21A (an example of a noise detection unit) that detects noise segments Nz included in the speech audio signals of the first and second speech segments. The similarity calculation unit 21D removes the noise segments Nz detected from the first and second speech segments and authenticates the speaker based on a comparison between the speech audio signals of the first and second speech segments from which the noise segments Nz have been removed and the registered speaker database DB. As a result, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using speech audio signals with speech features more suitable for user authentication processing by removing noise included in the speech audio signals, thereby improving user authentication accuracy.

[0144] Furthermore, the similarity calculation unit 21D in Embodiment 1 deletes either a first or second speech segment containing a noise segment Nz. If both the first and second speech segments are deleted, the speech segment detection unit 21A detects a third speech segment that is different from the first and second speech segments. If the similarity calculation unit 21D does not detect a noise segment Nz from the speech audio signal of the third speech segment by the speech segment detection unit 21A, it authenticates the speaker based on a comparison between the speech audio signal of the third speech segment and the registered speaker database DB. As a result, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using a speech audio signal with speech features more suitable for user authentication processing by removing the noise segment Nz contained in the speech audio signal, thereby improving user authentication accuracy.

[0145] Furthermore, the similarity calculation unit 21D in Embodiment 1 deletes either the first or second speech segment containing the noise segment Nz. If either the first or second speech segment is deleted, the speech segment detection unit 21A detects a third speech segment that is different from the first and second speech segments. If the noise detection unit does not detect a noise segment from the speech audio signal of the third speech segment, the similarity calculation unit 21D authenticates the speaker based on a comparison between the speech audio signal of the third speech segment and the other of the first or second speech segment that does not contain the noise segment Nz, and the registered speaker database DB. As a result, the authentication analysis device P1 according to Embodiment 1 can perform user authentication processing using a speech audio signal with speech features more suitable for user authentication processing by removing the speech segment containing noise, thereby improving user authentication accuracy.

[0146] Furthermore, the number of characters in Embodiment 1 is the number of moras, syllables, or phonemes. As a result, the authentication analysis device P1 according to Embodiment 1 can determine a speech audio signal or concatenated speech signal that has speech features more suitable for user authentication processing based on the number of moras, syllables, or phonemes. Therefore, the authentication analysis device P1 can improve user authentication accuracy even if there is variation in the speech features of the user included in each speech audio signal.

[0147] Although various embodiments have been described above with reference to the drawings, it goes without saying that this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the various embodiments described above can be combined arbitrarily without departing from the spirit of the invention.

[0148] This application is based on Japanese Patent Application No. 2021-157045 filed on September 27, 2021, and its contents are incorporated herein by reference. [Industrial applicability]

[0149] This disclosure is useful as an authentication device and authentication method for improving the accuracy of speaker voice authentication using spoken speech. [Explanation of Symbols]

[0150] 20 Communications Department 21 processors 21A Speech interval detection unit 21B Speech linking section 21C Feature Extraction Unit 21D Similarity calculation part 22 memory 22A ROM 22B RAM 23 Display I / F 100 Voice Recognition Systems DB Registered Speaker Database DP information display section Nz noise interval OP1 Operator-side call terminal P1 Authentication Analysis Device SC Authentication Results Screen US User UP1 User-side calling terminal

Claims

1. An acquisition unit that acquires the audio signal of the speaker's utterance, A detection unit that detects a first speech segment spoken by the speaker from the acquired audio signal, The system includes an authentication unit that authenticates the speaker based on a comparison between the audio signal of the first speech segment detected by the detection unit and a database. If the authentication unit determines that the speaker cannot be authenticated, the detection unit detects a second speech segment that is different from the first speech segment. The authentication unit authenticates the speaker based on a comparison between a concatenated speech signal, which is obtained by concatenating the speech signal of the first speech segment and the speech signal of the second speech segment, and the database. Authentication device.

2. The detection unit detects the first speech segment and the second speech segment, respectively, along the time series of the acquired audio signal. The authentication device according to claim 1.

3. The first utterance segment and the second utterance segment are two consecutive utterance segments detected by the detection unit. The authentication device according to claim 1.

4. The total length of the first speech segment and the second speech segment is equal to or greater than the first predetermined time. The authentication device according to claim 1.

5. The length of the first utterance and the second utterance are each at or above the second predetermined time. The authentication device according to claim 1.

6. The system further comprises a recognition unit that performs speech recognition of a first number of characters included in the first speech segment and a second number of characters included in the second speech segment, The total number of characters included in the first utterance section and the second utterance section is equal to or greater than the first predetermined number of characters. The authentication device according to claim 1.

7. The number of characters included in the first utterance section and the second utterance section is equal to or greater than the second predetermined number of characters. The authentication device according to claim 6.

8. The system further includes a processing unit that weights the first speech segment and the second speech segment based on the number of first and second characters recognized by the recognition unit, The processing unit calculates the total number of characters based on the first number of characters and the second number of characters, and performs weighting on the first utterance section based on the ratio of the first number of characters to the total number of characters, and weighting on the second utterance section based on the ratio of the second number of characters to the total number of characters. The authentication unit authenticates the speaker based on a comparison between the concatenated audio signal, which is obtained by concatenating the audio signal of the first utterance segment after weighting and the audio signal of the second utterance segment after weighting, and the database. The authentication device according to claim 6.

9. The system further comprises a conversion unit that converts the speech rate of the first speech segment and the speech rate of the second speech segment to a predetermined speech rate, The authentication unit authenticates the speaker based on a comparison between a concatenated audio signal, which is obtained by concatenating the audio signal of the first utterance section converted to a predetermined speech speed and the audio signal of the second utterance section converted to a predetermined speech speed, and the database. The authentication device according to claim 7.

10. The system further includes a noise detection unit that detects noise segments included in the audio signals of the first and second speech segments, The authentication unit removes the noise section detected from the first and second speech segments, and authenticates the speaker based on a concatenated speech signal obtained by concatenating the speech signal of the first speech segment from which the noise section has been removed and the speech signal of the second speech segment from which the noise section has been removed, and a comparison with the database. The authentication device according to claim 1.

11. The authentication unit deletes the first speech segment or the second speech segment that includes the noise segment. If both the first and second speech segments are deleted, the detection unit detects a third speech segment that is different from the first and second speech segments. If the noise detection unit does not detect the noise section from the audio signal of the third speech section, the authentication unit authenticates the speaker based on a comparison between the audio signal of the third speech section and the database. The authentication device according to claim 10.

12. The authentication unit deletes the first speech segment or the second speech segment that includes the noise segment. If either the first speech segment or the second speech segment is deleted, the detection unit detects a third speech segment that is different from the first speech segment and the second speech segment. If the noise detection unit does not detect the noise section from the audio signal of the third speech section, the authentication unit authenticates the speaker based on a comparison between the audio signal of the third speech section and the other of the first or second speech section that does not include the noise section, and the database. The authentication device according to claim 10.

13. The aforementioned number of characters is the number of moras, syllables, or phonemes. The authentication device according to any one of claims 6 to 8.

14. An authentication method performed by one or more computers, The audio signal of the speaker's utterance is acquired, From the acquired audio signal, the first speech segment spoken by the speaker is detected. Based on the matching of the detected audio signal of the first utterance segment with the database, the speaker is authenticated. If, based on the audio signal of the first utterance segment, it is determined that the speaker cannot be authenticated, a second utterance segment different from the first utterance segment is detected. The speaker is authenticated based on a concatenated speech signal obtained by concatenating the speech signals of the first speech segment and the speech signals of the second speech segment, and a comparison with the database. Authentication method.

Citation Information

Patent Citations

  • Voice processing device and program

    JP2009020459A

  • Authentication device, authentication method, and program

    JP2010109618A

  • Invalid voice input determination device, voice signal processing device, method, and program

    JP2016197200A

  • Voice processing device, voice processing method, and non-transitory computer readable medium having program stored thereon

    WO2020246041A1