Authentication device and authentication method
The authentication device and method provide real-time authentication status updates through voice recognition and database comparison, addressing the delay in existing systems and enhancing operator efficiency.
Patent Information
- Application Number
- JP2023565032
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-03
- Filing Date
- 2022-11-29
- Publication Date
- 2025-10-20
- Estimated Expiration
- 2042-11-29
AI Technical Summary
Existing operator identity verification systems at call centers do not display real-time progress of identity verification, leading to delayed initiation of the main conversation and reduced operator efficiency.
An authentication device and method that utilize voice recognition to authenticate speakers in real-time by comparing voice signals with a database, calculating reliability based on total time and sound types, and displaying authentication status updates.
Enables operators to check the authentication status of customer identity verification in real-time, improving work efficiency by allowing timely initiation of the main conversation.
Smart Images

Figure 0007756346000001 
Figure 0007756346000002 
Figure 0007756346000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an authentication device and an authentication method. [Background technology]
[0002] Patent Document 1 discloses an operator identity verification system that streamlines customer identity verification and other verification tasks at call centers. In this operator identity verification system, a voice recognition server recognizes the speech of the customer and the operator, outputs the speech, and stores the text of the speech recognition results along with date and time information. A keyword extraction unit in the analysis server then reads the text of the speech recognition results and extracts keywords for verification items included in the combination of the customer and the operator's speech from a predetermined verification item keyword list. A keyword matching unit in the analysis server then matches the extracted keywords for the verification items with the customer's basic member information stored in a member master database. If the two match, the server determines that verification of the verification item has been completed. When verification of all predetermined identity verification items has been completed, a notification of identity verification completion is displayed on the operator's terminal. The notification of identity verification completion is also sent to the customer's terminal. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2014-197140 Summary of the Invention [Problem to be solved by the invention]
[0004] In Patent Document 1, the operator's terminal does not display a message indicating that identity verification has been completed until all identity verification items prepared in advance for identity verification have been verified. Therefore, until all identity verification items have been verified, the progress of the authentication process to verify whether the operator is a legitimate customer (i.e., the customer) is not displayed in real time, and there is no way to grasp the current progress of the authentication process. As a result, the timing at which the operator begins the main topic of conversation, which begins after the customer's identity verification has been completed, is delayed, resulting in poor operator work efficiency.
[0005] The present disclosure has been devised in consideration of the above-described conventional situation, and aims to provide an authentication device and an authentication method that enable an operator to check the authentication status of customer identity verification in real time and help improve the operator's work efficiency. [Means for solving the problem]
[0006] The present disclosure provides an acquisition unit that acquires and detects a voice signal of a speaker's speech, and a method for authenticating whether the speaker is the person in question based on a comparison between the voice signal detected by the acquisition unit and a database. a first reliability based on the total time and a second reliability based on the number of sound types based on the total time and the number of sound types, based on the calculation results of the total time and the number of sound types and a predetermined determination criterion; and an authentication unit that performs authentication based on the authentication result of the authentication unit. , the first confidence level and the second confidence level; and a display interface that causes a terminal device to display an authentication status indicating whether the speaker is the person in question, wherein the display interface updates the display content of the authentication status every time the authentication status of the speaker by the authentication unit changes.
[0007] The present disclosure also provides a method for authentication performed by one or more computers, the method comprising: acquiring and detecting an audio signal of a speaker's speech; calculating a total time of the audio signal and the number of sound types included in the audio signal, and determining a first reliability based on the total time and a second reliability based on the number of sound types based on the calculation results of the total time and the number of sound types and a predetermined determination criterion; Based on the detected voice signal and a comparison with a database, it is determined whether the speaker is the person in question, and based on the result of the authentication, , the first confidence level and the second confidence level; The present invention provides an authentication method that displays an authentication status indicating whether the speaker is the person in question, and updates the displayed content of the authentication status every time the authentication status of the speaker changes.
[0008] These comprehensive or specific aspects may be realized as a system, an apparatus, a method, an integrated circuit, a computer program, or a recording medium, or may be realized as any combination of a system, an apparatus, a method, an integrated circuit, a computer program, and a recording medium. [Effects of the Invention]
[0009] According to the present disclosure, an operator can check the authentication status of customer identity verification in real time, thereby helping to improve the operator's work efficiency. [Brief explanation of the drawings]
[0010] [Figure 1] FIG. 1 is a diagram showing an example of a use case of the voice authentication system according to the first and second embodiments. [Figure 2] FIG. 1 is a block diagram showing an example of the internal configuration of an authentication analysis device according to a first embodiment. [Figure 3] FIG. 1 is a diagram showing a first relationship between a speech signal and a reliability according to the first embodiment; [Figure 4] FIG. 10 is a diagram showing a second relationship between a speech signal and reliability according to the first embodiment. [Figure 5] FIG. 10 is a diagram showing an example of starting authentication after pressing an authentication start button according to the first embodiment. [Figure 6A] FIG. 10 is a diagram showing the presence or absence of emotion in the audio signal according to the first embodiment. [Figure 6B] FIG. 10 shows processing of an audio signal depending on the presence or absence of emotion according to the first embodiment. [Figure 7] FIG. 10 is a diagram showing a process of deleting a repeated section of an audio signal according to the first embodiment; [Figure 8] FIG. 10 is a diagram showing a first example of a screen showing an authentication status according to the first embodiment; [Figure 9] FIG. 10 is a diagram showing a second example of a screen showing an authentication status according to the first embodiment; [Figure 10] 1 is a flowchart showing an example of an operation procedure of an authentication analysis device according to a first embodiment. [Figure 11] FIG. 10 is a block diagram showing an example of the internal configuration of an authentication analysis device according to a second embodiment. [Figure 12] FIG. 10 shows an example of a question according to the second embodiment. [Figure 13] FIG. 10 shows an example question displayed on an information terminal device according to the second embodiment. [Figure 14] FIG. 10 is a diagram showing the relationship between the number of phonemes and a threshold value according to the second embodiment. [Figure 15] 10 is a flowchart showing an example of an operation procedure of an authentication analysis device when displaying an example question immediately after starting authentication according to the second embodiment. [Figure 16] FIG. 10 is a diagram showing an example of a screen when the example question display function according to the second embodiment is turned off. [Figure 17] FIG. 10 is a diagram showing an example of a screen when the example question display function according to the second embodiment is on. [Figure 18] 10 is a flowchart showing an example of an operation procedure of an authentication analysis device when a sample question is displayed during authentication for personal identification according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, with reference to the drawings as appropriate, detailed descriptions will be given of embodiments specifically disclosing an authentication device and an authentication method in the present disclosure. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure and are not intended to limit the subject matter of the claims.
[0012] (Background to the first embodiment) Patent Document 1 discloses an operator identity verification system that streamlines customer identity verification and other verification tasks at call centers. In this operator identity verification system, a voice recognition server recognizes the speech of the customer and the operator, outputs the speech, and stores the text of the speech recognition results along with date and time information. A keyword extraction unit in the analysis server then reads the text of the speech recognition results and extracts keywords for verification items included in the combination of the customer and the operator's speech from a predetermined verification item keyword list. A keyword matching unit in the analysis server then matches the extracted keywords for the verification items with the customer's basic member information stored in a member master database. If the two match, the server determines that verification of the verification item has been completed. When verification of all predetermined identity verification items has been completed, a notification of identity verification completion is displayed on the operator's terminal. The notification of identity verification completion is also sent to the customer's terminal.
[0013] In Patent Document 1, the operator's terminal does not display a message indicating that identity verification has been completed until all identity verification items prepared in advance for identity verification have been verified. Therefore, until all identity verification items have been verified, the progress of the authentication process to verify whether the operator is a legitimate customer (i.e., the customer) is not displayed in real time, and there is no way to grasp the current progress of the authentication process. As a result, the timing at which the operator begins the main topic of conversation, which begins after the customer's identity verification has been completed, is delayed, resulting in poor operator work efficiency.
[0014] Therefore, in the following first embodiment, an example of an authentication device and authentication method will be described that enables an operator to check the authentication status of customer identity verification in real time and helps improve the operator's work efficiency.
[0015] (Embodiment 1) First, with reference to FIG. 1, a use case of the voice authentication system according to the first and second embodiments (see below) will be described. FIG. 1 is a diagram showing an example of a use case of the voice authentication system according to the first and second embodiments. The voice authentication system 100 acquires a voice signal or voice data of a person to be authenticated using voice (user US in the example shown in FIG. 1), and compares the acquired voice signal or voice data with multiple voice signals or voice data registered (stored) in advance in storage (registered speaker database DB in the example shown in FIG. 1). Based on the comparison result, the voice authentication system 100 evaluates the similarity between the voice signal or voice data collected from user US, who is the subject of voice authentication, and the voice signal or voice data registered in the storage, and authenticates user US based on the evaluated similarity.
[0016] The voice authentication system 100 according to the first embodiment includes at least an operator-side communication terminal OP1 as an example of a sound collection device, an authentication analysis device P1, a registered speaker database DB, and an information display terminal DP as an example of an output device. Note that the authentication analysis device P1 and the information display terminal DP may be integrally configured.
[0017] The voice authentication system 100 shown in Fig. 1 is used, as an example, to authenticate a speaker (user US) in a call center, and authenticates the user US using voice data obtained by collecting the speech of the user US while talking to an operator OP. The voice authentication system 100 shown in Fig. 1 further includes a user-side call terminal UP1 and a network NW. It goes without saying that the overall configuration of the voice authentication system 100 is not limited to the example shown in Fig. 1.
[0018] The user-side call terminal UP1 is connected to the operator-side call terminal OP1 via a network NW so as to be able to communicate wirelessly with the operator-side call terminal OP1. Note that the wireless communication here refers to network communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark).
[0019] The user-side communication terminal UP1 is configured by, for example, a notebook PC, a tablet terminal, a smartphone, a telephone, etc. The user-side communication terminal UP1 is a sound collection device equipped with a microphone (not shown), which collects the speech of the user US, converts it into a voice signal, and transmits this converted voice signal to the operator-side communication terminal OP1 via the network NW. The user-side communication terminal UP1 also acquires the voice signal of the speech of the operator OP transmitted from the operator-side communication terminal OP1, and outputs it from a speaker (not shown).
[0020] The network NW is, for example, an IP (Internet Protocol) network or a telephone network, and connects the user-side call terminal UP1 and the operator-side call terminal OP1 so that voice signals can be transmitted and received between them. Data is transmitted and received via wired or wireless communication.
[0021] The operator side communication terminal OP1 is connected to the user side communication terminal UP1 and the authentication analysis device P1 via wired or wireless communication so as to be able to transmit and receive data therebetween, and transmits and receives voice signals therebetween.
[0022] The operator-side communication terminal OP1 is configured, for example, by a notebook PC, a tablet terminal, a smartphone, a telephone, etc. The operator-side communication terminal OP1 acquires an audio signal based on the speech of the user US transmitted from the user-side communication terminal UP1 via the network NW, and transmits the audio signal to the authentication analysis device P1. When the operator-side communication terminal OP1 acquires an audio signal including the acquired speech of the user US and the speech of the operator OP, the operator-side communication terminal OP1 may separate the audio signal based on the speech of the user US from the audio signal based on the speech of the operator OP based on audio parameters such as the sound pressure level or frequency band of the audio signal of the operator-side communication terminal OP1. After the separation, the operator-side communication terminal OP1 extracts only the audio signal based on the speech of the user US and transmits it to the authentication analysis device P1.
[0023] The operator terminal OP1 may be communicably connected to each of a plurality of user terminals and simultaneously acquire voice signals from each of the plurality of user terminals. The operator terminal OP1 transmits the acquired voice signals to the authentication analysis device P1. This allows the voice authentication system 100 to simultaneously perform voice authentication processing and voice analysis processing for each of a plurality of users.
[0024] Furthermore, the operator-side communication terminal OP1 may simultaneously acquire a voice signal containing the respective speeches of multiple users. The operator-side communication terminal OP1 extracts a voice signal for each user from the voice signals of the multiple users acquired via the network NW and transmits the voice signal for each user to the authentication analysis device P1. In such a case, the operator-side communication terminal OP1 may analyze the voice signals of the multiple users and separate and extract the voice signals for each user based on voice parameters such as sound pressure level and frequency band. When the voice signals are collected by an array microphone or the like, the operator-side communication terminal OP1 may separate and extract the voice signals for each user based on the direction from which the spoken voices arrive. This allows the voice authentication system 100 to perform voice authentication processing and voice analysis processing for each of the multiple users, even if the voice signals are collected in an environment where multiple users are speaking simultaneously, such as a web conference.
[0025] An authentication analysis device P1, which is an example of an authentication device and a computer, is connected to an operator-side call terminal OP1, a registered speaker database DB, and an information display terminal DP so as to be able to transmit and receive data therebetween. Note that the authentication analysis device P1 may also be connected to the operator-side call terminal OP1, the registered speaker database DB, and the information display terminal DP via a network (not shown) so as to be able to communicate by wire or wirelessly.
[0026] The authentication analysis device P1 acquires the user US's voice signal transmitted from the operator-side communication terminal OP1 and performs voice analysis of the acquired voice signal, for example, by frequency, to extract individual speech features of the user US. The authentication analysis device P1 references the registered speaker database DB and compares the extracted speech features with the speech features of each of multiple users pre-registered in the registered speaker database DB to perform voice authentication of the user US. The authentication analysis device P1 may also compare the extracted speech features with the speech features of a specific user pre-registered in the registered speaker database DB instead of the speech features of each of multiple users pre-registered in the registered speaker database DB to perform voice authentication of the user US. The authentication analysis device P1 generates an authentication result screen SC including the user authentication result and transmits it to the information display terminal DP for output. The authentication result screen SC shown in FIG. 1 is merely an example and is not limited thereto. The authentication result screen SC shown in FIG. 1 includes, for example, a message indicating the user authentication result, such as "The voice of Taro Yamada has been matched." The authentication analysis device P1 may also perform voice authentication of the user US by comparing the voice signal of the user US with the voice signals of each of a plurality of users pre-registered in the registered speaker database DB. Note that the authentication analysis device P1 may also perform voice authentication of the user US by comparing the voice signal of the user US with the voice signal of a specific user pre-registered in the registered speaker database DB instead of the voice signals of each of a plurality of users pre-registered in the registered speaker database DB.
[0027] The registered speaker database DB, as an example of a database, is a so-called storage, and is configured using a storage medium such as a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The registered speaker database DB stores (registers) user information of each of a plurality of users in association with speech features. The user information here refers to information about a user, such as a user name, a user ID (Identification), or identification information assigned to each user. The registered speaker database DB may be configured integrally with the authentication analysis device P1.
[0028] The information display terminal DP is configured using, for example, an LCD (Liquid Crystal Display) or an organic EL (Electroluminescence) display. The information display terminal DP displays the authentication result screen SC transmitted from the authentication analysis device P1. The information display terminal DP may be configured integrally with the authentication analysis device P1.
[0029] 1, the user-side communication terminal UP1 collects the user US's speech COM12 "This is Yamada Taro" and speech COM14 "This is 123245678," converts them into audio signals, and transmits them to the operator-side communication terminal OP1. The operator-side communication terminal OP1 transmits audio signals based on the user US's speech COM12 and COM14 transmitted from the user-side communication terminal UP1 to the authentication analysis device P1.
[0030] When the operator-side communication terminal OP1 acquires audio signals that have been collected from the operator OP's speech COM11 "Please tell me your name," the speech COM13 "Please tell me your membership number," and the user US's speech COM12 and COM14, it separates and removes the audio signals based on the operator OP's speech COM11 and COM13, extracts only the audio signals based on the user US's speech COM12 and COM14, and transmits them to the authentication analysis device P1. This allows the authentication analysis device P1 to improve the accuracy of user authentication by using only the audio signals of the person who is the target of voice authentication.
[0031] Next, an example of the internal configuration of the authentication analysis device will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the internal configuration of the authentication analysis device according to embodiment 1. The authentication analysis device P1 includes at least a communication unit 20, a processor 21, and a memory 22.
[0032] The communication unit 20 is connected to the operator-side communication terminal OP1 and the registered speaker database DB so as to enable data communication between them. The communication unit 20 outputs a voice signal transmitted from the operator-side communication terminal OP1 to the processor 21. Note that the acquisition unit is not limited to the communication unit 20, and may be, for example, a microphone of the operator-side communication terminal OP1 configured integrally with the authentication analysis device P1.
[0033] The processor 21 is configured using a semiconductor chip on which at least one of electronic devices such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), a GPU (Graphical Processing Unit), an FPGA (Field Programmable Gate Array), etc. The processor 21 functions as a controller that manages the overall operation of the authentication analysis device P1, and performs control processing for managing the operation of each part of the authentication analysis device P1, data input / output processing between each part of the authentication analysis device P1, data arithmetic processing, and data storage processing.
[0034] Processor 21 realizes the functions of speech section detection unit 21A, speech connection unit 21B, feature extraction unit 21C, similarity calculation unit 21D, reliability calculation unit 21E, and sound analysis unit 21J by using programs and data stored in ROM (Read Only Memory) 22A of memory 22. Processor 21 uses RAM (Random Access Memory) 22B of memory 22 during operation, and temporarily stores data or information generated or acquired by processor 21 and each unit in RAM 22B of memory 22.
[0035] The speech section detection unit 21A, as an example of an acquisition unit, acquires an audio signal of an utterance, analyzes the acquired audio signal, and detects an utterance section in which the user US is speaking. The speech section detection unit 21A outputs an audio signal corresponding to each utterance section detected from the audio signal (hereinafter referred to as an "utterance audio signal") to the speech connection unit 21B or the feature extraction unit 21C. The speech section detection unit 21A may also temporarily store the utterance audio signal of each utterance section in the RAM 22B of the memory 22.
[0036] When the speech section detection unit 21A detects two or more speech sections of the same person (user US) from the audio signal, the speech connection unit 21B, as an example of an authentication unit, connects the speech audio signals of these speech sections. The speech connection unit 21B may calculate the total number of seconds of the connected audio signals. The speech connection unit 21B outputs the connected speech audio signal (hereinafter referred to as the "connected audio signal") to the feature extraction unit 21C. The user authentication method will be described later.
[0037] The feature extraction unit 21C, as an example of an authentication unit, analyzes the characteristics of an individual's voice, for example, for each frequency, using one or more speech voice signals extracted by the speech period detection unit 21A to extract speech features. Note that the feature extraction unit 21C may also extract speech features of the concatenated voice signal output from the speech concatenation unit 21B. The feature extraction unit 21C associates the extracted speech features with the speech voice signal or concatenated voice signal from which the speech features were extracted, and outputs them to the similarity calculation unit 21D or temporarily stores them in RAM 22B of the memory 22.
[0038] The similarity calculation unit 21D, which is an example of an authentication unit, acquires the speech feature of the speech voice signal or the concatenated voice signal output from the feature extraction unit 21C. The similarity calculation unit 21D refers to the registered speaker database DB and calculates the similarity between the speech feature of each of the multiple users registered in the registered speaker database DB and the acquired concatenated speech feature. Based on the calculated similarity, the similarity calculation unit 21D identifies the user corresponding to the speech voice signal or the concatenated voice signal (i.e., the voice signal transmitted from the user-side call terminal UP1) and performs user authentication.
[0039] If the similarity calculation unit 21D determines that the user has been identified as a result of the user authentication, it generates an authentication result screen SC including information about the identified user (i.e., the authentication result) and outputs it to the information display terminal DP via the display I / F (Interface) 23.
[0040] If the similarity calculation unit 21D determines that the calculated similarity is less than a predetermined value, it may determine that user authentication is not possible, and generate and output a control command requesting the speech connection unit 21B to connect speech voice signals. Furthermore, if an upper limit is set for the number of times user authentication is attempted for the same person (user US), and the similarity calculation unit 21D determines that the number of times user authentication is not possible is equal to or exceeds the upper limit, it may generate an authentication result screen (not shown) notifying that user authentication is not possible, and output it to the information display terminal DP.
[0041] The sound analysis unit 21J, as an example of an authentication unit, acquires one or more speech signals or connected speech signals extracted by the speech segment detection unit 21A. The sound analysis unit 21J analyzes the sounds of the acquired speech signals or connected speech signals (hereinafter referred to as "speech sounds"). For example, if the speech signal is "Yamada Taro desu," the corresponding sounds are "ya ma da ta ro u de su." The sound analysis unit 21J may calculate the number of analyzed speech sounds, generate a calculation result screen (not shown), and output it to the information display terminal DP. In the first embodiment, the number of types of speech sounds is defined as one sound, such as "ya," which is a combination of one consonant and one vowel.
[0042] The reliability calculation unit 21E, as an example of an authentication unit, acquires a connected voice signal connected by the speech connection unit 21B. The reliability calculation unit 21E analyzes the reliability of the acquired connected voice signal. The reliability calculation unit 21E may acquire a speech voice signal extracted by the speech segment detection unit 21A and analyze the reliability. The reliability may be, for example, but is not limited to, the reliability of the total number of seconds calculated by the speech connection unit 21B or the reliability of the number of speech sound types calculated by the sound analysis unit 21J. The reliability calculation unit 21E may calculate the reliability based on a determination criterion predetermined by the user.
[0043] As a result, processor 21 authenticates whether the speaker is the person in question by comparing the speech sound signal detected by speech period detection unit 21A with a registered speaker database DB in which the speech signals of each of a plurality of speakers are registered. Processor 21 also calculates the total time of the speech sound signal and the number of speech sound types included in the speech sound signal, and determines a first reliability based on the total time and a second reliability based on the number of speech sound types based on the calculation results of the total time and the number of speech sound types and a predetermined determination criterion.
[0044] The memory 22 includes at least a ROM 22A that stores programs that define the various processes performed by the processor 21 and data used during execution of the programs, and a RAM 22B that serves as a work memory used when the various processes performed by the processor 21 are executed. The ROM 22A stores programs that define the various processes performed by the processor 21 and data used during execution of the programs. The RAM 22B temporarily stores data or information generated or acquired by the processor 21 (for example, speech audio signals before concatenation, concatenated speech signals after concatenation, speech features corresponding to each speech section before or after concatenation, etc.).
[0045] The display I / F 23 connects the processor 21 and the information display terminal DP so that data communication is possible between them, and outputs to the information display terminal DP an authentication result screen SC generated by the similarity calculation unit 21D of the processor 21. The display I / F 23 causes the information display terminal DP to display an authentication status indicating whether or not the speaker is the correct person based on the authentication result of the processor 21.
[0046] The emotion classifier 24 is connected to the processor 21 so as to be able to communicate data with it, and can be realized using, for example, artificial intelligence (AI). The emotion classifier 24 is configured using, for example, a processor such as a GPU (Graphical Processing Unit) that can execute various processes using artificial intelligence. The emotion classifier 24 detects the intensity of a speaker's emotion when speaking based on the speech audio signal detected by the speech segment detection unit 21A. The emotion classifier 24 may obtain a concatenated audio signal concatenated by the speech concatenation unit 21B and detect the intensity of a speaker's emotion when speaking, or may detect the intensity of a speaker's emotion when speaking using features of the audio signal extracted by the feature extraction unit 21C. To detect the intensity of an emotion, the emotion classifier 24 analyzes, for example, but is not limited to, the volume, pitch (frequency), and accent of the audio signal.
[0047] Next, a first relationship between a voice signal and reliability will be described with reference to Fig. 3. Fig. 3 is a diagram showing the first relationship between a voice signal and reliability according to Embodiment 1. It goes without saying that the user's utterance contents shown in Figs. 3 and 4 are merely examples and are not limited to these.
[0048] 3 is temporarily stored in, for example, the memory 22, and indicates the relationship between the speech sound, the total number of seconds, the number of speech sound types, and the reliability of the user's speech content. Note that the elements of the reliability of the user's speech content are not limited to the total number of seconds and the number of speech types, and the number of reliability elements is not limited to two.
[0049] In the example shown in FIG. 3, the higher reliability of the total number of seconds or the number of speech sound types is determined as the final reliability of the user's speech content. That is, the reliability calculation unit 21E determines the higher of the first reliability based on the total time or the second reliability based on the number of speech sound types as the reliability corresponding to the speech audio signal. For example, the reliability of the total number of seconds is determined as "low" for less than 10 seconds, "medium" for 10 to 15 seconds, and "high" for 15 seconds or more. The reliability of the number of speech sound types is determined as "low" for less than 10 sounds, "medium" for 10 to 15 sounds, and "high" for 15 sounds or more. Note that the reliability determination criteria are merely examples and are not limited to these. Reliability can be expressed using characters such as "low," "medium," and "high," as well as percentages, gauges, and bar graphs.
[0050] The speech sounds of the first utterance C1 "Yamada ta ro u de su" are "ya ma da ta ro u de su". The total length of the first utterance C1 is 5 seconds, and the number of speech sound types is 8. In this case, the reliability of the total length of time is "low", and the reliability of the number of speech sound types is "low". As a result, since the reliability of both the total length of time and the number of speech types are "low", the reliability of the first utterance C1 is "low".
[0051] The second utterance C2 is "Yamada Taro desu Yamada Jiro to Shiro desu" and consists of the speech sounds "ya ma da ta ro u de su ya ma da ji ro u to shi ro u de su". The total length of the second utterance C2 is 10 seconds, and the number of speech sounds is 11. In this case, the reliability of the total length of time is "medium", and the reliability of the number of speech sounds is also "medium". As a result, since the reliability of both the total length of time and the number of speech sounds is "medium", the reliability of the second utterance C2 is "medium".
[0052] The speech sounds of the third utterance C3 "Yamada ta ro de su i chi ni sa N shi go ro ku na na de su" are "ya ma da ta ro u de su i chi ni sa N shi go ro ku na na de su." The total length of the third utterance C3 is 10 seconds, and the number of speech sounds is 18. In this case, the reliability of the total length of time is "medium," and the reliability of the number of speech sounds is "high." In the example shown in FIG. 3, the higher of the reliability of the total length of time or the reliability of the number of speech sounds is determined as the reliability corresponding to the user's utterance. Therefore, the reliability of the number of speech sounds, which is more reliable, "high," is determined as the reliability of the third utterance C3.
[0053] As a result, when either the reliability of the total number of seconds or the number of speech sound types reaches a reliability threshold, the reliability of the user's speech content is determined to be above the threshold, thereby shortening the time required to complete authentication of the speaker's identity.
[0054] Next, a second relationship between a speech signal and a reliability will be described with reference to Fig. 4. Fig. 4 is a diagram showing the second relationship between a speech signal and a reliability according to the first embodiment.
[0055] 4 is temporarily stored in, for example, memory 22, and indicates the relationship between the speech sound, total number of seconds, number of speech sound types, and reliability of the user's speech content. Note that the elements of the reliability of the user's speech content are not limited to the total number of seconds and the number of speech sound types, and the number of elements of reliability is not limited to two.
[0056] In the example shown in FIG. 4, the reliability of the total number of seconds or the reliability of the number of speech sound types, whichever is lower, is determined as the final reliability determination of the user's speech content. That is, the reliability calculation unit 21E determines the lower of the first reliability based on the total time or the second reliability based on the number of speech sound types as the reliability corresponding to the speech audio signal. For example, the reliability determination criteria for the total number of seconds are set as "low" for less than 8 seconds, "medium" for 8 to 10 seconds, and "high" for 10 seconds or more. The reliability determination criteria for the number of speech sound types are set as "low" for less than 9 sounds, "medium" for 10 to 15 sounds, and "high" for 15 sounds or more. It goes without saying that the reliability determination criteria are merely examples and are not limited to these. The example shown in FIG. 4 differs from the example shown in FIG. 3 only in the final reliability determination method, and therefore a description of the parts that overlap with the example shown in FIG. 3 will be omitted.
[0057] In the example shown in FIG. 4, a bar graph is used to represent the reliability. The reliability Ba1 of the first total number of seconds will be described as an example. The reliability Ba1 is a rectangular bar graph that is long from side to side, and as the reliability increases relative to a predetermined reference value, the meter increases continuously from left to right. When the reliability becomes "high" at the predetermined reference value, the meter reaches the right end. Note that this is just an example, and a vertically long rectangular bar graph may also be used to represent the reliability. This allows the reliability to be treated as a continuous parameter.
[0058] The total number of seconds of the first utterance content C1 is 5 seconds, and the number of types of speech sounds is 8. In this case, the reliability of the total number of seconds is "low", and the reliability of the number of types of speech sounds is "low". The reliability of the total number of seconds is lower than the reliability Ba1 and the reliability Ba2. Therefore, the reliability of the first utterance content C1 is the reliability Ba1, "low".
[0059] The total number of seconds for the second utterance content C2 is 10 seconds, and the number of speech sound types is 11. In this case, the reliability of the total number of seconds is "high," and the reliability of the number of speech sound types is "medium." For reliability Ba3, the meter reaches the right end because the reliability of the total number of seconds is "high." Because the number of speech sound types for the second utterance content C2 is three more than the number of speech sound types for the first utterance content C1, the meter for reliability Ba4 increases compared to reliability Ba2. Looking at reliability Ba3 and reliability Ba4, it can be seen that reliability Ba4 is less reliable, and the reliability of the second utterance content C2 becomes "medium" for reliability Ba4.
[0060] The total number of seconds for the third utterance content C3 is 10 seconds, and the number of speech sound types is 18. In this case, the reliability of the total number of seconds is "high", and the reliability of the number of speech sound types is "high". For reliability Ba5, the meter reaches the right end as the reliability of the total number of seconds becomes "high". For reliability Ba6, the meter reaches the right end as the reliability of the number of speech types becomes "high". Based on reliability Ba5 and reliability Ba6, the reliability of the third utterance content C3 becomes "high" for reliability Ba5 or reliability Ba6.
[0061] As a result, when a lower reliability level reaches the reliability threshold, the reliability of the user's speech content is determined to be equal to or greater than the threshold, thereby increasing the reliability of authentication for speaker identity confirmation.
[0062] Next, the timing of starting authentication will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example in which authentication starts after the authentication start button according to the first embodiment is pressed.
[0063] In the example shown in FIG. 5, first, the user US and the operator OP converse over the following utterances: the operator OP's speech COM15 "This is the XX call center," the user US's speech COM16 "I would like to do XX," the operator's speech COM17 "First, we will perform identity authentication," and the user US's speech COM18 "Yes." The operator OP then presses the authentication start button UI displayed on the information display terminal DP, and the authentication analysis device P1 starts collecting audio from the speech signal after the authentication start button UI is pressed. The user US and the operator OP then converse over the following utterances: the operator OP's speech COM11 "Please tell me your name," the user US's speech COM12 "I am Yamada Taro," the operator OP's speech COM13 "Please tell me your membership number," and the user US's speech COM14 "I am 12345678." In this case, after the operator OP presses the authentication start button UI displayed on the information display terminal DP, audio signals based on the speech sounds COM12 and COM14 of the user US are transmitted to the authentication analysis device P1. In other words, when the processor 21 receives a signal indicating that the authentication start button UI displayed on the information display terminal DP has been pressed, it starts authentication from the speech sound signals input after receiving the signal.
[0064] The authentication analysis device P1 performs authentication based on the acquired speech voice signal and displays the authentication result on the authentication result screen SC on the information display terminal DP. In the example shown in Figure 5, when identity verification is completed, the authentication result screen SC displays "The voice matches that of Taro Yamada."
[0065] The authentication analysis device P1 can intentionally exclude unnecessary speech signals of the user US from the authentication for identity verification based on the operation of the operator OP. This allows the authentication analysis device P1 to use only the necessary speech signals of the user US for identity verification, thereby improving authentication accuracy.
[0066] Next, emotion of a concatenated audio signal will be described with reference to Fig. 6A. Fig. 6A is a diagram showing the presence or absence of emotion in the audio signal according to Embodiment 1. Note that the audio signal in Fig. 6A may be a speech audio signal.
[0067] In the graph shown in Figure 6A, the horizontal axis represents time and the vertical axis represents the intensity of emotion. On the horizontal axis, time progresses as you move to the right, and on the vertical axis, the intensity of emotion increases as you move up.
[0068] Emotion waveform Wa1 is a waveform that represents the strength of an emotion identified by emotion classifier 24. Processor 21 determines that emotion waveform Wa1 is present when it is equal to or greater than a predetermined threshold, and that emotion is absent when it is below the threshold. That is, in the example shown in Fig. 6A, section S1 of emotion waveform Wa1 is present, and other sections are absent.
[0069] Next, processing of an audio signal depending on the presence or absence of emotion will be described with reference to Fig. 6B. Fig. 6B is a diagram showing processing of an audio signal depending on the presence or absence of emotion according to Embodiment 1. Note that the audio signal in Fig. 6B may be a speech audio signal or a concatenated audio signal.
[0070] The audio signal waveforms Sig2, Sig3, and Sig4 are concatenated audio signals, and are waveforms that represent the strength of the audio signals.
[0071] The processor 21 determines the presence or absence of emotion based on the intensity of emotion in the audio signal waveform Sig2 detected by the emotion classifier 24. As a result, it determines that sections S2 and S3 of the audio signal waveform Sig2 contain no emotion, and section S4 contains an emotion. The processor 21 uses only sections S2 and S3 of the audio signal waveform Sig2 that are determined to contain no emotion for authentication, and does not use section S4 that is determined to contain emotion for authentication. The authentication analysis device P1 deletes the audio signal in section S4 and generates a single concatenated audio signal by concatenating the audio signal waveform Sig3 in section S2 and the audio signal waveform Sig4 in section S3. In other words, the processor 21 determines whether the detection result of the intensity of emotion is equal to or greater than a predetermined threshold, and deletes the audio signal in the audio section where the intensity of emotion is equal to or greater than the predetermined threshold.
[0072] This allows the authentication analysis device P1 to delete voice signal sections that are not suitable for authentication of personal identification due to heightened emotions, thereby improving the accuracy of authentication.
[0073] Next, a process for deleting a repetitive section of an audio signal will be described with reference to Fig. 7. Fig. 7 is a diagram showing a process for deleting a repetitive section of an audio signal according to the first embodiment.
[0074] The speech signal waveform Sig5 is a concatenated speech signal "Yes, it's Yamada. It's Yamada Taro. Yes, please." When the processor 21 analyzes the speech signal waveform Sig5, it finds that "Yamada," "desu," and "yes" appear repeatedly within the speech signal waveform Sig5. The processor 21 determines to use the section S5 "Yes, it's Yamada," the section S6 "Taro," and the section S7 "Thank you very much" for authentication. On the other hand, the processor 21 determines not to use the section S8 "Yamada" and the section S9 "desu yes," which contain overlapping content, for authentication. The processor 21 deletes sections S8 and S9, and generates a single concatenated speech signal by concatenating the speech signal waveform Sig6 of section S5, the speech signal waveform Sig7 of section S6, and the speech signal waveform Sig8 of section S7. That is, the processor 21 recognizes the speech signal, detects speech sections with overlapping speech content from the speech recognition results of the speech signal, and deletes the speech signal of the detected overlapping speech sections.
[0075] This allows the authentication analysis device P1 to delete the audio signal in sections where the speech content overlaps, thereby improving the accuracy of authentication.
[0076] Next, examples of screens showing authentication status will be described with reference to Fig. 8 and Fig. 9. Fig. 8 is a diagram showing a first example of a screen showing authentication status according to embodiment 1, and Fig. 9 is a diagram showing a second example of a screen showing authentication status according to embodiment 1. It should be noted that these screen examples are merely examples, and needless to say, the present invention is not limited to these.
[0077] The display DP1 is an example of a screen displayed on the information display terminal DP by the display I / F 23. The content displayed on the display DP1 includes at least candidate authentication results for the speaker and the reliability of the authentication result for personal identification.
[0078] Message Msg1 displays information about the person who is currently closest to user US among the user information stored in the registered speaker database DB during identity verification. For example, the content displayed in message Msg1 may be a person's photo, name, gender, address, or phone number in the candidate information field MN1. Note that the content displayed in message Msg1 is merely an example and is not limited to these.
[0079] The authentication result candidate field MN2 displays candidates for the authentication result of identity verification. The authentication result candidate field MN2 may display the names of the candidates together with the probability that each candidate is the user US. The probability that each candidate is the user US may be displayed as a bar meter as shown in the authentication result candidate field MN2 in FIG. 8, or as a percentage. The candidates with the highest probability of being the user US may be displayed in order from top to bottom, in kana alphabet order, or in alphabetical order, or the order may be arbitrarily set by the operator OP.
[0080] The audio signal display field MN3 displays the waveform of the concatenated audio signal. The audio signal waveform Sig9 is a concatenated speech audio signal, and is a waveform that indicates the strength of the audio signal. The audio signal display field MN3 displays the sections used for authentication of the audio signal waveform Sig9 and the sections that are not used so that they can be clearly seen. For example, in the example shown in Figure 8, the sections S10, S11, S12, and S13 used for authentication of the audio signal waveform Sig9 are displayed with different background colors. This makes it possible to visualize the sections with emotions and the sections where the speech content is repeated in the audio signal waveform Sig9. In other words, the display I / F 23 displays the determination result of the processor 21 regarding the presence or absence of emotions on the information display terminal DP. The operator OP may select unnecessary audio sections based on the determination result displayed by the display I / F 23 and delete the audio signal of the selected audio sections.
[0081] The authentication result reliability meter MN4 displays the reliability of the number of speech phonemes and the total number of seconds of the speech signal waveform Sig9 on a meter.
[0082] The button BT1 is an authentication start / stop button, which allows the operator OP to start authentication from the voice signal uttered after pressing the button BT1.
[0083] FIG. 9 shows an example of a screen when authentication has progressed further than the authentication status shown in FIG.
[0084] The authentication result candidate field MN2 displays real-time candidates and the probability that each candidate is user US as identity verification progresses. The authentication result candidate field MN2 shown in Fig. 9 shows an increased probability that user US is "Yamada Taro" compared to the authentication result candidate field MN2 shown in Fig. 8. This allows the operator OP to know the candidates for the authentication result of real-time identity verification.
[0085] The voice signal waveform Sig10 displayed in the voice signal display field MN3 is a connected voice signal when the dialogue between the operator OP and the user US has progressed more than in the voice signal waveform Sig9. In contrast to the voice signal waveform Sig9, the voice signal in the section S14 of the voice signal waveform Sig10 is added as the voice to be used for authentication.
[0086] The authentication result reliability meter MN4 is a meter displaying the reliability of the number of speech sound types and the total number of seconds of the audio signal waveform Sig10. The reliability of the number of speech sound types and the total number of seconds of the audio signal waveform Sig10 increases compared to the example shown in Figure 8 because the audio signal of section S14 of the audio signal waveform Sig9 has been added as the audio to be used for authentication. This allows the operator OP to know the reliability of the number of speech phoneme types and the total number of seconds in real time during identity verification authentication.
[0087] As in the examples of FIGS. 8 and 9, the display I / F 23 updates the display content of the authentication status every time the authentication status of the speaker by the processor 21 changes.
[0088] As a result, the authentication analysis device P1 displays the authentication result of the identity verification by the processor 21 on the information display terminal DP in real time, allowing the operator OP to check the authentication status of the identity verification in real time, thereby improving the work efficiency of the operator OP.
[0089] Next, an example of the operation procedure of the authentication analysis device will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of the operation procedure of the authentication analysis device according to the first embodiment.
[0090] The communication unit 20 in the authentication analysis device P1 acquires the voice signal (or voice data) transmitted from the operator-side communication terminal OP1 (St11).
[0091] The display I / F 23 in the authentication analysis device P1 acquires a signal indicating whether the authentication start button displayed on the information display terminal DP has been pressed (St12). If the display I / F 23 has not acquired a signal indicating that the authentication start button has been pressed (St12, NO), the process returns to step St11. If the display I / F 23 has acquired a signal indicating that the authentication start button has been pressed (St12, YES), the display I / F 23 outputs the audio signal acquired by the communication unit 20 in the process of step St11 to the processor 21.
[0092] When the display I / F 23 acquires a signal indicating that the authentication start button has been pressed in the process of step St12, the processor 21 starts authentication of the user US who is the voice authentication target of the acquired voice signal (St13).
[0093] The speech period detection unit 21A in the processor 21 detects a speech period from the acquired audio signal (St14).
[0094] The speech section detection unit 21A stores information about the detected speech section (e.g., start time and end time of the speech section, number of characters, number of speech sound types, signal length (speech audio length, number of seconds of speech, etc.), speech speed before or after speech speed conversion, etc.) in memory 22 (St15).
[0095] The speech section detection unit 21A selects one or more speech voice signals to be used for user authentication based on the currently set user authentication processing method (St16). Although not shown in Fig. 10, if the authentication analysis device P1 determines that there is no speech voice signal to be used for user authentication based on the currently set user authentication processing method, the authentication analysis device P1 may return to the processing of step St14 and detect a new speech section. The speech section detection unit 21A outputs the selected speech voice signals to the speech connection unit 21B.
[0096] The utterance connection unit 21B executes a voice connection process to connect each of the one or more selected utterance voice signals, and generates a connected voice signal (St17). The utterance connection unit 21B outputs the generated connected voice signal to the reliability calculation unit 21E.
[0097] The reliability calculation unit 21E calculates the reliability using the concatenated audio signal generated in the process of step St17 (St18). For example, the reliability calculated in the process of step St18 is the reliability of the total number of seconds of the concatenated audio signal and the reliability of the number of types of speech sounds. Note that the reliability calculated in the process of step St18 is not limited to these.
[0098] The display I / F 23 displays or updates the display of the reliability calculated in the process of step St18 on the information display terminal DP (St19).
[0099] The reliability calculation unit 21E determines whether the reliability calculated in the process of step St18 is equal to or greater than a predetermined threshold (St20). If the reliability is less than the threshold in the process of step St20 (NO in St20), the authentication analysis device P1 determines whether to continue the authentication for personal identification (St21). When the reliability calculation unit 21E determines, for example, that the current authentication count is less than a predetermined upper limit of the authentication count (YES in St21), the process of the processor 21 returns to the process of step St14. When the reliability calculation unit 21E determines, for example, that the current authentication count is equal to or greater than a predetermined upper limit of the authentication count (NO in St21), the processor 21 determines that the user authentication has failed based on the acquired voice signal (St22). The display I / F 23 generates an authentication result screen notifying the user of the failure of user authentication and outputs the screen to the information display terminal DP. The information display terminal DP outputs (displays) the authentication result screen transmitted from the authentication analysis device P1.
[0100] On the other hand, if the reliability is equal to or greater than the threshold value in the processing of step St20 (St20, YES), the reliability calculation unit 21E outputs the concatenated speech signal to the feature extraction unit 21C. The feature extraction unit 21C extracts the speech features of the individual user US from the concatenated speech signal output from the reliability calculation unit 21E (St23). The feature extraction unit 21C outputs the extracted speech features of the individual user US to the similarity calculation unit 21D.
[0101] The similarity calculation unit 21D refers to the speech features of each of the multiple users registered in the registered speaker database DB, and calculates the similarity between the speech features of the individual user US output from the feature extraction unit 21C and the speech features of each of the multiple users registered in the registered speaker database DB (St24).
[0102] The similarity calculation unit 21D determines whether or not there is a user whose calculated similarity is equal to or greater than a threshold value among the multiple users registered in the registered speaker database DB (St25).
[0103] In the process of step St25, if it is determined that there is a user whose calculated similarity is equal to or greater than a threshold among the multiple users registered in the registered speaker database DB (St25, YES), the similarity calculation unit 21D determines that this user is the user US of the voice signal (St26). Note that, if it is determined that there are multiple users whose similarity is equal to or greater than a threshold, the similarity calculation unit 21D may determine that the user with the highest similarity is the user US of the voice signal.
[0104] If the similarity calculation unit 21D determines that the user has been identified, it outputs information about the identified user (i.e., the authentication result) to the display I / F 23, and the display I / F 23 generates an authentication result screen SC based on the information output by the similarity calculation unit 21D and outputs it to the information display terminal DP (St27).
[0105] On the other hand, if the similarity calculation unit 21D determines in the processing of step St25 that there is no user among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than the threshold value (St25, NO), it determines whether to continue the identity verification authentication by, for example, determining whether the current number of user authentication processes is equal to or greater than the set upper limit number (St21).
[0106] In the process of step St21, when determining whether to continue the authentication for personal identification, for example, if the similarity calculation unit 21D determines that the current authentication count is equal to or greater than a predetermined upper limit of the authentication count (NO in St21), it determines that the user authentication has failed based on the acquired voice signal (St22). The display I / F 23 generates an authentication result screen notifying the user that the user authentication has failed and outputs it to the information display terminal DP. The information display terminal DP outputs (displays) the authentication result screen transmitted from the authentication analysis device P1.
[0107] When the similarity calculation unit 21D determines that the current number of authentication attempts is less than the predetermined upper limit of the number of authentication attempts (St21, YES), the similarity calculation unit 21D returns to the processing of step St14.
[0108] As described above, the authentication analysis device P1 according to the first embodiment can perform user authentication processing using a speech signal that is more suitable for user authentication processing by a predetermined user authentication processing method. As a result, the authentication analysis device P1 according to the first embodiment can improve the accuracy of user authentication.
[0109] As described above, the authentication analysis device P1 according to the first embodiment comprises a speech section detection unit 21A that acquires and detects the audio signal of the speaker's speech, a processor 21 that authenticates whether the speaker is the person in question based on a comparison with the registered speaker database DB, and a display I / F 23 that displays the authentication status indicating whether the speaker is the person in question based on the authentication result of the processor 21 on the information display terminal DP, and the display I / F 23 updates the display content of the authentication status every time the speaker's authentication status by the processor 21 changes.
[0110] As a result, the authentication analysis device P1 displays the authentication result of the personal identification by the processor 21 on the information display terminal DP in real time, allowing the operator OP to check the authentication status of the personal identification in real time, thereby improving the work efficiency of the operator OP.
[0111] Furthermore, the processor 21 according to the first embodiment calculates the total time of the audio signal and the number of sounds included in the audio signal, and determines a first reliability based on the total time and a second reliability based on the number of sound types based on the calculation results of the total time and the number of sound types and a predetermined determination criterion. As a result, the reliability of the authentication result of the speaker's identity confirmation is calculated and the reliability is notified to the operator OP in real time, so that the operator OP can predict the timing of completion of speaker authentication and the work efficiency of the operator OP can be improved.
[0112] Furthermore, the processor 21 according to the first embodiment determines the higher of the first reliability and the second reliability as the reliability corresponding to the speech signal. As a result, the reliability determination can be completed when either the first reliability or the second reliability satisfies a predetermined determination criterion, thereby shortening the time required to complete speaker authentication.
[0113] Furthermore, the processor 21 according to the first embodiment determines the lower of the first reliability and the second reliability as the reliability corresponding to the voice signal. This allows the reliability determination to be completed when both the first reliability and the second reliability satisfy a predetermined determination criterion, thereby improving authentication accuracy.
[0114] Furthermore, when the processor 21 according to the first embodiment receives a signal indicating that the authentication start button displayed on the information display terminal DP has been pressed, it starts authentication from the voice signal input after the signal is received. This allows the start of authentication for confirming the identity of the speaker to be initiated by an operation of the operator OP, so that the operator OP can inform the user US that authentication will begin before starting it. Furthermore, if the operator OP determines that authentication is not necessary, it is possible to not perform authentication.
[0115] The authentication analysis device P1 according to the first embodiment further includes an emotion classifier 24 that detects the intensity of the emotion of the speaker when speaking based on the audio signal, and the processor 21 determines whether the detected result of the emotion intensity is equal to or greater than a predetermined threshold, and deletes the audio signal of the audio section where the emotion intensity is equal to or greater than the predetermined threshold. This allows the authentication analysis device P1 to improve authentication accuracy by detecting and deleting audio sections that are not suitable for identity verification.
[0116] The authentication analysis device P1 according to the first embodiment further includes an emotion classifier 24 that detects the intensity of the emotion of the speaker when speaking based on the audio signal, and the processor 21 determines whether the detection result of the emotion intensity is equal to or greater than a predetermined threshold, and the display I / F 23 displays the determination result on the information display terminal DP and deletes the audio signal of the audio section selected by a user operation of the determination result displayed by the display I / F 23. This allows the operator OP to arbitrarily delete audio sections that are not suitable for authentication of the detected identity, thereby improving authentication accuracy.
[0117] Furthermore, the processor 21 according to the first embodiment recognizes the speech signal, detects speech sections in which speech content overlaps from the speech recognition result of the speech signal, and deletes the detected speech signal of the overlapping speech section. This allows the authentication analysis device P1 to delete the speech sections in which speech content overlaps from the speech speech signal and the concatenated speech signal, thereby making authentication for identity verification more efficient.
[0118] Furthermore, the display content of the authentication status according to the first embodiment includes at least authentication result candidates of the speaker and the reliability of the authentication result, which allows the operator OP to check the authentication status of the identity verification in real time, thereby improving the work efficiency of the operator OP.
[0119] (Background to the second embodiment) In Patent Document 1, the operator's customer identity verification task involves simply displaying predetermined verification items on the operator terminal. This identity verification task can result in the total time of the acquired voice signal being short, or the phonemes of the acquired voice signal being biased. This results in poor authentication accuracy.
[0120] Therefore, in the following embodiment 2, an example of an authentication device and authentication method that allows an operator to perform authentication of customer identity confirmation with high accuracy will be described. In the following description, the same components as those in embodiment 1 will be denoted by the same reference numerals, and their description will be omitted.
[0121] (Embodiment 2) An example of the internal configuration of an authentication analysis device will be described with reference to Fig. 11. Fig. 11 is a block diagram showing an example of the internal configuration of an authentication analysis device according to embodiment 2. The authentication analysis device P2 includes at least a communication unit 20, a processor 21H, and a memory 22H. In embodiment 2, the memory 22H further includes an example question sentence data storage unit 22C, as compared with embodiment 1.
[0122] Processor 21H uses programs and data stored in ROM 22A of memory 22H to realize the functions of speech section detection unit 21A, utterance connection unit 21B, feature extraction unit 21C, similarity calculation unit 21D, phoneme analysis unit 21F, and example sentence selection unit 21G. Processor 21H uses RAM 22B of memory 22H during operation, and temporarily stores data or information generated or acquired by processor 21H and each unit in RAM 22B of memory 22H.
[0123] The example sentence selection unit 21G, as an example of an authentication unit, selects an example question to be displayed on the information display terminal DP from among a plurality of example questions stored in the example question data storage unit 22C. The example sentence selection unit 21G selects an appropriate example question and displays it on the information display terminal DP in order to improve the accuracy of authentication for identity confirmation. The example sentence selection unit 21G may select an example question immediately after the start of authentication and display it on the information display terminal DP. Furthermore, the example sentence selection unit 21G may select an example question and display it on the information display terminal DP when the authentication progresses and the similarity calculation unit 21D determines that the similarity between the speech voice signal or the connected voice signal is equal to or less than a threshold. Furthermore, the example sentence selection unit 21G may select an example question and display it on the information display terminal DP as the authentication progresses based on the analysis results of the phoneme analysis unit 21F.
[0124] The example question data storage unit 22C, which is an example of an authentication unit, stores data on example questions selected by the example sentence selection unit 21G and displayed on the information display terminal DP. The example question data storage unit 22C stores a plurality of example questions for acquiring a voice signal used for speaker authentication by the processor 21H. The example question data storage unit 22C may be provided in the memory 22H, or may be external to the authentication analysis device P2 and connected to the authentication analysis device P2 so as to be able to communicate data with it.
[0125] The phoneme analysis unit 21F, which is an example of an authentication unit, extracts phonemes contained in the voice signal of the speaker detected by the speech period detection unit 21A. Here, the definition of the phonemes calculated by the phoneme analysis unit 21F according to the second embodiment will be described. For example, if the speech voice signal is "Yamada Tarou desu," the corresponding phoneme is "yamadataroudesu." In other words, in the second embodiment, the number of speech phonemes is defined such that each consonant and vowel, such as "y" and "a," count as one phoneme.
[0126] Next, an example question will be described with reference to Fig. 12. Fig. 12 is a diagram showing an example question according to the second embodiment.
[0127] Table TBL3 shows examples of question items other than "name" and "membership number." The question items shown in table TBL3 include "address," "telephone number," "date of birth," "password for telephone procedures," and "kana characters." It goes without saying that these are just examples of question items and are not limited to these.
[0128] "Address" in table TBL3 is an example question with first priority because it can obtain more spoken phonemes than "Name" and "Membership Number."
[0129] "Telephone number" in table TBL3 has fewer phonemes than "address", but "date of birth" has many good phonemes, so it is an example question with second priority.
[0130] "Date of birth" in table TBL3 has fewer spoken phonemes than "telephone number," and is therefore an example question with priority 3.
[0131] The "Password for telephone procedures" in table TBL3 is an example question issued in advance by the company, which contains a password containing phonemes that are not included in the personal information such as "Name," "Membership number," "Address," "Telephone number," and "Date of birth."
[0132] The "kana characters" in table TBL3 are example questions in kana characters that elicit additional spoken phonemes that were not obtained from the user US's speech signal in response to the example questions of "address," "telephone number," and "date of birth." The above-mentioned spoken phonemes are analyzed by the phoneme analysis unit 21F. For example, if a spoken phoneme in the "ka" row cannot be obtained, the example question becomes, "For identity authentication, could you please say 'ka kaki ku ke ko'?"; if a spoken phoneme in the "ta" row cannot be obtained, the example question becomes, "For identity authentication, could you please say 'tachi tsu te to'?"
[0133] In this way, the example questions are questions that prompt the speaker to respond with at least one of an address, a telephone number, a date of birth, and a password or kana characters that include phonemes that are not included in the speaker's personal information.
[0134] Next, an example question displayed on the information terminal device will be described with reference to Fig. 13. Fig. 13 is a diagram showing an example question displayed on the information terminal device according to the second embodiment.
[0135] The question screen Msg2 is an example of a screen of example questions displayed on the information display terminal DP. The example questions related to the question screen Msg2 are selected by the example sentence selection unit 21G. Note that the question screen Msg2 is not limited to this.
[0136] The question screen Msg2 displays, as an example question for priority 1, "May I have your registered address?". The question screen Msg2 displays, as an example question for priority 2, "May I have your registered phone number?". The question screen Msg2 displays, as an example question for priority 3, "May I have your date of birth?". The question screen Msg2 displays, as an example question for priority 4, "For identity verification, may I have your password for telephone procedures?". The question screen Msg2 displays, as an example question for priority 5, "For identity verification, could you please say 'ka kaki ku ke ko' (Ka row)?".
[0137] The question screen Msg2 may display multiple example questions at once, or may display only the example question with the highest priority.
[0138] In this way, the display I / F 23 causes the example question selected by the example sentence selection unit 21G to be displayed on the information display terminal DP.
[0139] Next, the relationship between the number of each speech phoneme calculated by the phoneme analysis unit and the threshold value will be described with reference to Fig. 14. Fig. 14 is a diagram showing the relationship between the number of each phoneme and the threshold value according to the second embodiment.
[0140] Graph Gr1 is a bar graph showing the number of speech phonemes calculated by the phoneme analysis unit 21F. Graph Gr1 shows a predetermined threshold and speech phonemes below the threshold. In graph Gr1, speech phoneme L1 “k,” speech phoneme L2 “t,” speech phoneme L3 “r,” and speech phoneme L4 “j” are speech phonemes with speech phoneme counts below the threshold. The example sentence selection unit 21G may select an example question based on the speech phonemes L1, L2, L3, and L4. For example, the example sentence selection unit 21G may select an example question from the example question data storage unit 22C, from which at least one of speech phonemes L1, L2, L3, and L4 can be collected from the speech audio signal when the user speaks. Furthermore, graph Gr1 may be arranged in alphabetical order based on the number of speech phonemes for each speech phoneme calculated by phoneme analysis unit 21F, or in order of the number of speech phonemes. Although graph Gr1 displays only speech phonemes included in the speech voice signal or the concatenated voice signal, it may also display speech phonemes that are not included.
[0141] As a result, the example sentence selection unit 21G selects example question sentences that encourage speech including speech phonemes that are not included in the speech sound signal or the connected sound signal, based on the speech phonemes extracted by the phoneme analysis unit 21F. Also, the example sentence selection unit 21G selects example question sentences that include speech phonemes whose number of speech phonemes is less than a predetermined threshold, based on the number of speech phonemes for each of the speech phonemes extracted by the phoneme analysis unit 21F.
[0142] Next, an operation procedure of the authentication analysis device in the case where an example question sentence is displayed immediately after the start of authentication will be described with reference to Fig. 15. Fig. 15 is a flowchart showing an example of an operation procedure of the authentication analysis device in the case where an example question sentence is displayed immediately after the start of authentication according to the second embodiment.
[0143] The communication unit 20 in the authentication analysis device P2 acquires the voice signal (or voice data) transmitted from the operator-side communication terminal OP1 (St31).
[0144] The display I / F 23 in the authentication analysis device P2 acquires a signal indicating whether the authentication start button displayed on the information display terminal DP has been pressed (St32). If the display I / F 23 has not acquired a signal indicating that the authentication start button has been pressed (St32, NO), the process returns to step St31. If the display I / F 23 has acquired a signal indicating that the authentication start button has been pressed (St32, YES), the display I / F 23 outputs the audio signal acquired by the communication unit 20 in the process of step St31 to the processor 21H.
[0145] When the display I / F 23 acquires a signal indicating that the authentication start button has been pressed in the process of step St32, the processor 21H starts authentication of the user US who is the voice authentication target of the acquired voice signal (St33).
[0146] The example sentence selection unit 21G acquires example question sentences from the example question sentence data storage unit 22C and selects example question sentences to be displayed on the information display terminal DP. The example sentence selection unit 21G transmits a signal including the content of the selected example question sentence to the display I / F 23. When the display I / F 23 acquires the signal including the content of the selected example question sentence, it displays the example question sentence selected by the example sentence selection unit 21G immediately after the start of authentication (St34).
[0147] The speech period detector 21A in the processor 21H detects a speech period from the acquired audio signal (St14).
[0148] The speech section detection unit 21A stores information about the detected speech section (e.g., start time and end time of the speech section, number of characters, number of speech phonemes, signal length (speech audio length, number of seconds of speech, etc.), speech speed before or after speech speed conversion, etc.) in memory 22H (St36).
[0149] The speech period detection unit 21A selects one or more speech voice signals to be used for user authentication based on the currently set user authentication processing method (St37). Although not shown in Fig. 10, if the authentication analysis device P2 determines that there is no speech voice signal to be used for user authentication based on the currently set user authentication processing method, the authentication analysis device P2 may return to the processing of step St35 and detect a new speech period. The speech period detection unit 21A outputs the selected speech voice signals to the phoneme analysis unit 21F.
[0150] The phoneme analysis unit 21F executes a process of analyzing the speech phonemes of the speech voice signal selected in the process of step St37 (St38). The phoneme analysis unit 21F outputs the analyzed speech voice signal to the speech connection unit 21B.
[0151] The speech connection unit 21B executes a speech connection process to connect the one or more selected speech signals to generate a connected speech signal (St39). The speech connection unit 21B outputs the generated connected speech signal to the similarity calculation unit 21D.
[0152] The similarity calculation unit 21D calculates the similarity between the speech signal of the speaker's response to the example question and the speech signal registered in the registered speaker database DB (St40). Although not shown in FIG. 15, the utterance connection unit 21B may output the connected speech signal generated in the processing of step St39 to the feature extraction unit 21C. That is, the similarity calculation unit 21D may refer to the speech features of each of the multiple users and calculate the similarity between the speech feature of the individual user US output from the feature extraction unit 21C and the speech feature of each of the multiple users registered in the registered speaker database DB. The similarity calculation unit 21D may calculate the similarity with the speech feature of a specific user registered in the registered speaker database DB instead of the speech feature of each of the multiple users registered in the registered speaker database DB.
[0153] The similarity calculation unit 21D transmits a signal including the calculated similarity to the display I / F 23. When the display I / F 23 receives the signal including the similarity, it causes the information display terminal DP to display the result of the calculated similarity (St41).
[0154] The similarity calculation unit 21D determines whether or not there is a user whose calculated similarity is equal to or greater than a predetermined threshold value among the multiple users registered in the registered speaker database DB (St42).
[0155] In the process of step St42, if it is determined that there is a user whose calculated similarity is equal to or greater than a threshold among the multiple users registered in the registered speaker database DB (St42, YES), the similarity calculation unit 21D determines that this user is the user US of the voice signal (St45). Note that, if it is determined that there are multiple users whose similarity is equal to or greater than a threshold, the similarity calculation unit 21D may determine that the user with the highest similarity is the user US of the voice signal.
[0156] If the similarity calculation unit 21D determines that the user has been identified, it outputs information about the identified user (i.e., the authentication result) to the display I / F 23, and the display I / F 23 generates an authentication result screen SC based on the information output by the similarity calculation unit 21D and outputs it to the information display terminal DP (St46).
[0157] On the other hand, if the similarity calculation unit 21D determines in the processing of step St42 that there is no user among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than the threshold value (St42, NO), the similarity calculation unit 21D determines whether to continue the identity verification authentication (St43).
[0158] In the process of step St43, when determining whether to continue the authentication for personal identification, for example, if the similarity calculation unit 21D determines that the current authentication count is equal to or greater than a predetermined upper limit of the authentication count (NO in St43), it determines that the user authentication has failed based on the acquired voice signal (St44). The display I / F 23 generates an authentication result screen notifying the user that the user authentication has failed and outputs it to the information display terminal DP. The information display terminal DP outputs (displays) the authentication result screen transmitted from the authentication analysis device P2 (St46).
[0159] When determining whether to continue the authentication for personal identification, for example, if it is determined that the current number of authentication times is less than the predetermined upper limit of the number of authentication times (St43, YES), the similarity calculation unit 21D returns to the processing of step St34.
[0160] Next, an example of a screen when the example question display function is turned off will be described with reference to Fig. 16. Fig. 16 is a diagram showing an example of a screen when the example question display function according to the second embodiment is turned off.
[0161] The display DP2 is an example of a screen showing the authentication status displayed on the operator OP.
[0162] Information IF1 displays personal information of the candidate of the authentication result. The personal information may be name, telephone number, address, or membership number. However, the personal information is not limited to these. Information IF1 may also display the first candidate of the authentication result, or multiple candidates. Information IF2 displays a photograph of the candidate of the authentication result.
[0163] The authentication result candidate field MN5 displays candidates for the authentication result of identity verification. The authentication result candidate field MN5 may display the names of the candidates together with the probability that each candidate is user US. The probability that each candidate is user US may be a bar meter as shown in the authentication result candidate field MN5 of FIG. 8, or may be expressed as a percentage. The authentication result candidate field MN5 may display the candidates with the highest probability of being user US in ascending order, or may display the candidates in kana alphabet order or alphabetical order, or the order of the candidates may be arbitrarily set by the operator OP.
[0164] The example question display field MN6 displays the example question selected by the example sentence selection unit 21G. In the example of Fig. 16, the example question display function is turned off, so the example question is not displayed in the example question display field MN6.
[0165] The voice signal display field MN7 displays the waveform of the connected voice signal in real time. When the phoneme analysis unit 21F is analyzing the spoken phonemes of the connected voice signal, the voice signal display field MN7 may display "Phonetic analysis in progress."
[0166] Button BT2 is an authentication start / stop button. By pressing button BT2, the operator OP can start and stop the personal identification authentication.
[0167] The button BT3 is a button for instructing whether or not the example question display function is to be turned on / off. By pressing the button BT3, the operator OP can operate whether or not example questions are to be displayed in the example question display field MN6.
[0168] Information IF3 displays the number of speech phonemes, speech length (i.e., total time), and number of speech sections in real time.
[0169] Next, an example of a screen when the example question display function is on will be described with reference to Fig. 17. Fig. 17 is a diagram showing an example of a screen when the example question display function according to embodiment 2 is on. Note that explanations of parts that overlap with Fig. 16 will be omitted.
[0170] When the operator OP presses the button BT3 to turn on the example question display function, the example question selected by the example sentence selection unit 21G is displayed in the example question display field MN6. For example, the example question display field MN6 displays "Example question: For identity verification, could you please say 'ka kaki ku ke ko' (Ka row)?"
[0171] Note that the screen examples shown in FIGS. 16 and 17 are merely examples, and the present invention is not limited to these.
[0172] Next, an operation procedure of the authentication analysis device when displaying an example question during identity verification authentication will be described with reference to Fig. 18. Fig. 18 is a flowchart showing an example operation procedure of the authentication analysis device when displaying an example question during identity verification authentication according to the second embodiment.
[0173] The communication unit 20 in the authentication analysis device P2 acquires the voice signal (or voice data) transmitted from the operator-side communication terminal OP1 (St51).
[0174] The display I / F 23 in the authentication analysis device P2 acquires a signal indicating whether the authentication start button displayed on the information display terminal DP has been pressed (St52). If the display I / F 23 has not acquired a signal indicating that the authentication start button has been pressed (St52, NO), the process returns to step St51. If the display I / F 23 has acquired a signal indicating that the authentication start button has been pressed (St52, YES), the display I / F 23 outputs the audio signal acquired by the communication unit 20 in the process of step St51 to the processor 21.
[0175] When the display I / F 23 acquires a signal indicating that the authentication start button has been pressed in the process of step St32, the processor 21H starts authentication of the user US who is the voice authentication target of the acquired voice signal (St53).
[0176] The speech period detector 21A in the processor 21H detects a speech period from the acquired audio signal (St54).
[0177] The speech section detection unit 21A stores information about the detected speech section (e.g., start time and end time of the speech section, number of characters, number of speech phonemes, signal length (speech audio length, number of seconds of speech, etc.), speech speed before or after speech speed conversion, etc.) in memory 22H (St55).
[0178] The speech period detection unit 21A selects one or more speech voice signals to be used for user authentication based on the currently set user authentication processing method (St56). Although not shown in Fig. 10, if the authentication analysis device P2 determines that there is no speech voice signal to be used for user authentication based on the currently set user authentication processing method, the authentication analysis device P2 may return to the processing of step St54 and detect a new speech period. The speech period detection unit 21A outputs the selected speech voice signals to the phoneme analysis unit 21F.
[0179] The phoneme analysis unit 21F executes a process of analyzing the speech phonemes of the speech voice signal selected in the process of step St56 (St57). The phoneme analysis unit 21F outputs the analyzed speech voice signal to the speech connection unit 21B.
[0180] The speech connection unit 21B executes a speech connection process to connect the one or more selected speech signals to generate a connected speech signal (St58). The speech connection unit 21B outputs the generated connected speech signal to the similarity calculation unit 21D.
[0181] The similarity calculation unit 21D calculates the similarity between the speech voice signal of the speaker's response to the example question sentence and the speech voice signal registered in the registered speaker database DB (St59). Note that, although not shown in Fig. 16, the utterance connection unit 21B may output the connected speech signal generated in the processing of step St39 to the feature extraction unit 21C. In other words, the similarity calculation unit 21D may refer to the speech features of each of the multiple users and calculate the similarity between the speech features of the individual user US output from the feature extraction unit 21C and the speech features of each of the multiple users registered in the registered speaker database DB.
[0182] The similarity calculation unit 21D transmits a signal including the calculated similarity to the display I / F 23. When the display I / F 23 acquires the signal including the similarity, it causes the information display terminal DP to display the result of the calculated similarity (St60).
[0183] The similarity calculation unit 21D determines whether or not there is a user whose calculated similarity is equal to or greater than a predetermined threshold value among a plurality of users registered in the registered speaker database DB (St61).
[0184] In the process of step St61, if it is determined that there is a user whose calculated similarity is equal to or greater than a threshold among the multiple users registered in the registered speaker database DB (St61, YES), the similarity calculation unit 21D determines that this user is the user US of the voice signal (St62). Note that, if it is determined that there are multiple users whose similarity is equal to or greater than a threshold, the similarity calculation unit 21D may determine that the user with the highest similarity is the user US of the voice signal.
[0185] If the similarity calculation unit 21D determines that the user has been identified, it outputs information about the identified user (i.e., the authentication result) to the display I / F 23, and the display I / F 23 generates an authentication result screen SC based on the information output by the similarity calculation unit 21D and outputs it to the information display terminal DP (St63).
[0186] On the other hand, if the similarity calculation unit 21D determines in the processing of step St61 that there is no user among the multiple users registered in the registered speaker database DB whose calculated similarity is equal to or greater than the threshold value (St61, NO), the similarity calculation unit 21D continues the authentication for identity verification (St64).
[0187] In the process of step St21, when determining whether to continue the authentication for personal identification, for example, if the similarity calculation unit 21D determines that the current authentication count is equal to or greater than a predetermined upper limit of the authentication count (NO in St64), it determines that the user authentication has failed based on the acquired voice signal (St65). The display I / F 23 generates an authentication result screen notifying the user that the user authentication has failed and outputs it to the information display terminal DP. The information display terminal DP outputs (displays) the authentication result screen transmitted from the authentication analysis device P2 (St63).
[0188] When the similarity calculation unit 21D determines whether to continue the authentication for personal identification, for example, if it determines that the current authentication count is less than a predetermined upper limit of the authentication count (YES in St64), the similarity calculation unit 21D outputs the determination result to the example sentence selection unit 21G. The example sentence selection unit 21G determines whether to display an example question sentence (St66). The determination of whether to display an example question sentence may be made by the similarity calculation unit 21D, or may be made by the example sentence selection unit 21G based on a predetermined threshold value related to the number of speech phonemes or similarity, or may be made based on whether the operator OP has pressed a button displayed on the information display terminal DP for indicating whether to display an example question sentence.
[0189] If the example sentence selection unit 21G determines that display of the example question sentence is not necessary (St66, not necessary), the process returns to step St54. If the example sentence selection unit 21G determines that display of the example question sentence is necessary (St66, necessary), the process outputs the determination result to the display I / F 23. The display I / F 23 displays the example question sentence on the information display terminal DP (St67), and the process returns to step St54.
[0190] 18, the processor 21H calculates the similarity between the speech signal acquired after the start of authentication and the speech signal registered in the registered speaker database DB. If the similarity is equal to or less than a predetermined threshold, the processor 21H determines whether or not it is necessary to display an example question. If the processor 21H obtains a determination result that it is necessary to display an example question, the display I / F 23 causes the information display terminal DP to display the example question selected by the example sentence selection unit 21G.
[0191] As described above, the authentication analysis device P2 according to the second embodiment includes a speech section detection unit 21A that acquires and detects the audio signal of the speaker's speech, a processor 21H that authenticates whether the speaker is the person in question based on a comparison with the registered speaker database DB, a sample question data storage unit 22C that stores a plurality of sample questions for acquiring the audio signal used for speaker authentication by the processor 21H, a display interface 23 that displays the sample questions for the speaker on the information display terminal DP, and a sample question selection unit 21G that selects a sample question to be displayed on the information display terminal DP from the plurality of sample questions stored in the sample question data storage unit 22C.
[0192] This allows the authentication analysis device P2 to select a sample question sentence to display on the information display terminal DP in order to acquire the voice signal required for authenticating the identity of the speaker, thereby enabling the operator OP to authenticate the identity of the customer with high accuracy.
[0193] The authentication analysis device P2 according to the second embodiment further includes a phoneme analysis unit 21F that extracts speech phonemes contained in the speaker's voice signal detected by the speech segment detection unit 21A. The example sentence selection unit 21G selects an example question that prompts speech containing speech phonemes not contained in the voice signal, based on the speech phonemes extracted by the phoneme analysis unit 21F. This allows the authentication analysis device P2 to select an example question for eliciting uncollected speech phonemes. This allows the operator OP to efficiently authenticate the speaker's identity.
[0194] Furthermore, the display I / F 23 according to the second embodiment displays the example question selected by the example sentence selection unit 21G on the information display terminal DP. This allows the operator OP to ask the customer an example question for extracting uncollected speech phonemes, thereby enabling highly accurate authentication of the customer's identity.
[0195] Furthermore, the display I / F 23 according to the second embodiment displays the example question selected by the example sentence selection unit 21G on the information display terminal DP immediately after the start of authentication. The processor 21H calculates the similarity between the speech signal of the speech uttered by the speaker in response to the example question and the speech signal registered in the registered speaker database DB, and authenticates the speaker if the similarity is equal to or greater than a predetermined threshold. This allows the operator OP to smoothly ask the speaker an example question that elicits the speech phonemes required for authenticating the speaker's identity. This allows the operator OP to efficiently authenticate the speaker's identity.
[0196] Furthermore, the processor 21H according to the second embodiment calculates the similarity between the voice signal acquired after the start of authentication and the voice signal registered in the registered speaker database DB, and if the similarity is equal to or less than a predetermined threshold, determines whether or not it is necessary to display an example question. When the processor 21H obtains a determination result that it is necessary to display an example question, the display I / F 23 causes the information display terminal DP to display the example question selected by the example sentence selection unit 21G. This causes the authentication analysis device P2 to display, on the information display terminal DP, an example question for extracting speech phonemes necessary for identity verification. This allows the operator OP to perform identity verification with high accuracy.
[0197] Furthermore, the phoneme analysis unit 21F according to the second embodiment calculates the number of each extracted phoneme. The example sentence selection unit 21G selects an example question sentence including speech phonemes whose number of speech phonemes is less than a predetermined threshold. This allows the authentication analysis device P2 to display, on the information display terminal DP, an example question sentence for extracting speech phonemes whose number of speech phonemes is less than the threshold from among the collected speech phonemes. This allows the operator OP to perform authentication of the customer's identity with high accuracy.
[0198] Furthermore, when the processor 21H according to the second embodiment acquires a signal indicating that the authentication start button displayed on the information display terminal DP has been pressed, it starts authentication from the voice signal input after acquiring the signal. This allows the start of authentication for confirming the identity of the speaker to be initiated by an operation of the operator OP, so that the operator OP can inform the user US that authentication will begin before starting it. Furthermore, when the operator OP determines that authentication is not necessary, it is possible to not perform authentication.
[0199] The example questions according to the second embodiment are questions that require the speaker to respond with at least one of the following: address, telephone number, date of birth, a password containing phonemes not included in the speaker's personal information, or kana characters. This allows the authentication analysis device P2 to efficiently acquire the speech signal used for identity verification. This allows the operator OP to perform identity verification with high accuracy and also allows authentication to be performed efficiently in a short time.
[0200] Although the embodiments have been described above with reference to the accompanying drawings, the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also fall within the technical scope of the present disclosure. Furthermore, the components in the above-described embodiments may be combined in any manner without departing from the spirit of the invention.
[0201] This application is based on a Japanese patent application (Patent Application No. 2021-197229) filed on December 3, 2021, the contents of which are incorporated by reference into this application. [Industrial Applicability]
[0202] The technology disclosed herein is useful for providing an authentication device and authentication method that enables an operator to check the authentication status of a customer's identity verification in real time, thereby helping to improve the operator's work efficiency and performing identity verification with high accuracy. [Explanation of symbols]
[0203] NW Network UP1 User side call terminal OP1 Operator side call terminal US users OP Operator COM11,COM12,COM13,COM14,COM15,COM16,COM17,COM18 Speech audio P1, P2 authentication analysis device DB Registered speaker database DP Information Display Terminal SC authentication result screen 20 Communications Department 21,21H processor 21A Speech activity detector 21B Speech connector 21C Feature Extraction Unit 21D Similarity calculation part 21E Reliability calculation unit 21F Phoneme analysis department 21G Example sentence selection section 22,22H memory 21J Sound Analysis Department 22A ROM 22B RAM 22C Question example data storage section 23 Display I / F 24 Emotion Classifier Tables TBL1, TBL2, TBL3 C1, C2, C3 Speech content Ba1,Ba2,Ba3,Ba4,Ba5,Ba6 Reliability UI authentication start button Wa1 emotion waveform S1,S2,S3,S4,S5,S6,S7,S8,S9,S10,S11,S12,S13,S14 section Sig2, Sig3, Sig4, Sig5, Sig6, Sig7, Sig8, Sig9, Sig10 Audio signal waveform DP1, DP2 displays MN1 Candidate information column MN2, MN5 authentication result candidate column MN3,MN7 Audio signal display field MN4 Authentication Result Reliability Meter MN6 Example question display field Msg1 message Msg2 Question Screen BT1,BT2,BT3 Gr1 graph L1,L2,L3,L4 phoneme IF1, IF2, IF3 information
Claims
1. an acquisition unit that acquires and detects a speech signal of a speaker; an authentication unit that authenticates whether the speaker is the person in question based on the voice signal detected by the acquisition unit and a comparison with a database, calculates the total time of the voice signal and the number of sound types included in the voice signal, and determines a first reliability based on the total time and a second reliability based on the number of sound types based on the calculation results of the total time and the number of sound types and a predetermined determination criterion; a display interface that displays, on a terminal device, an authentication status indicating whether the speaker is the person in question, including the first reliability and the second reliability, based on the authentication result of the authentication unit; the display interface updates the display content of the authentication status every time the authentication status of the speaker by the authentication unit changes. Authentication device.
2. the authentication unit determines the higher of the first reliability and the second reliability as the reliability corresponding to the voice signal; The authentication device according to claim 1 .
3. the authentication unit determines the lower reliability of the first reliability or the second reliability as the reliability corresponding to the voice signal; The authentication device according to claim 1 .
4. when the authentication unit receives a signal indicating that an authentication start button displayed on the terminal device has been pressed, the authentication unit starts authentication from the voice signal input after receiving the signal. The authentication device according to any one of claims 1 to 3.
5. an emotion classifier that detects the intensity of the emotion of the speaker when speaking based on the audio signal; the authentication unit determines whether the detection result of the intensity of emotion is equal to or greater than a predetermined threshold, and deletes the audio signal of a voice section in which the intensity of emotion is equal to or greater than the predetermined threshold. The authentication device according to claim 1 .
6. an emotion classifier that detects the intensity of the emotion of the speaker when speaking based on the audio signal; the authentication unit determines whether the detection result of the intensity of the emotion is equal to or greater than a predetermined threshold; the display interface causes the terminal device to display the result of the determination; deleting the speech signal of the speech section selected by a user operation on the determination result displayed on the terminal device by the display interface; The authentication device according to claim 1 .
7. the authentication unit performs speech recognition on the speech signal, detects a speech section in which speech content overlaps from the speech recognition result of the speech signal, and deletes the speech signal of the detected overlapping speech section. The authentication device according to claim 1 .
8. the display content includes at least authentication result candidates of the speaker and authentication result reliability of the authentication; The authentication device according to claim 1 .
9. 1. A method of authentication performed by one or more computers, comprising: Acquiring and detecting an audio signal of a speaker's speech; Calculating the total duration of the audio signal and the number of types of sounds included in the audio signal; determining a first reliability based on the total time and a second reliability based on the number of sound types based on a predetermined determination criterion; authenticating the identity of the speaker based on the detected voice signal and a comparison with a database; displaying an authentication status indicating whether the speaker is a real person or not, including the first reliability and the second reliability based on the authentication result; updating the display content of the authentication status every time the authentication status of the speaker changes; Authentication method.
Citation Information
Patent Citations
Talker check system
JP2005300958A
Conference support system and conference support method
JP2010055307A
Customer identity verification support system for operator, and method therein
JP2014197140A
Contact history management system
JP2017054356A
Speaker verification system
US20120284026A1