Registration information update method and registration information update device
The method and device improve voice authentication accuracy and security by updating registration information based on voice recognition, addressing issues of content mismatch and impersonation in voice authentication systems.
Patent Information
- Application Number
- JP2024057072
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-10-10
AI Technical Summary
Existing voice authentication methods face challenges such as decreased accuracy due to differences in spoken content during registration and authentication, difficulty in determining what to speak, and vulnerability to impersonation, especially when user information changes over time.
A method and device that update registration information by selecting and correcting user information based on voice recognition of spoken content, ensuring accurate and secure authentication by managing unchanging and changing user data.
Enhances voice authentication accuracy and security by automatically updating user information, preventing impersonation, and reducing authentication failures due to changed user details.
Smart Images

Figure 2025154199000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a registration information updating method and a registration information updating device. [Background technology]
[0002] Patent Document 1 discloses an information processing device that acquires a registered first voice of a first user identified in a first authentication using an authentication method other than an authentication method using voice, acquires a second voice uttered by a second user during an operation, compares the characteristics of the first voice with the characteristics of the second voice, and performs a second authentication to confirm that the second user is the first user identified in the first authentication, and if the second authentication is successful, executes processing according to the operation associated with the second voice. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-7620 Summary of the Invention [Problem to be solved by the invention]
[0004] In the past, when performing voice authentication using spontaneous speech, the content (phonemes) spoken during registration differed from the content (phonemes) spoken during authentication, and the features extracted from the speech during registration and the speech during authentication were not similar, which could result in a decrease in authentication accuracy. Furthermore, when users were asked to speak freely, it was difficult for them to determine what content to speak. One approach to address this problem was to have the user speak a lot in order to collect the phonemes necessary for voice authentication, but this method had the problem of taking a long time for authentication.
[0005] Furthermore, voice authentication makes it easy for a third party to impersonate a user by recording the user's voice. Therefore, there is a demand for a voice authentication method that can prevent impersonation by a third party. Patent Document 1 discloses a method for preventing impersonation by a third party by combining another authentication method with voice authentication.
[0006] Meanwhile, as a voice authentication method that prevents impersonation by a third party without using a physical authentication tool (an IC card in Patent Document 1), in addition to authentication based on voice features, there is a voice authentication method that registers user information (e.g., the user's name, date of birth, address, or telephone number) in advance, recognizes the voice during voice authentication, and determines whether the voice recognition result matches the registered user information.
[0007] However, in voice authentication, when the content of an utterance is compared with pre-registered user information (for example, the user's name, date of birth, address, or telephone number), there is a possibility that the actual user information may have changed between the time of registration and the time of authentication. In other words, even if a user utters their current user information (for example, their address), the voice authentication may fail because the registered user information (for example, their address) is different.
[0008] The present disclosure has been devised in consideration of the above-described conventional situation, and aims to provide a registration information update method and registration information update device that assist in the management of user information used for voice authentication and suppress a decline in voice authentication accuracy. [Means for solving the problem]
[0009] The present disclosure provides a method for updating registration information performed by at least one computer, which selects first registration information to be used for speaker authentication from a database in which multiple registration information related to registered speakers is registered, acquires a first speech audio signal obtained by having the speaker speak regarding the first registration information, performs voice recognition on the first speech audio signal to acquire the first speech content of the speaker, and if it is determined that the speaker information included in the first speech content does not match the first registration information, selects second registration information different from the first registration information, acquires a second speech audio signal obtained by having the speaker speak regarding the second registration information, performs voice recognition on the second speech audio signal to acquire the second speech content of the speaker, and if it is determined that the speaker information included in the second speech content matches the second registration information, corrects the first registration information registered in the database to the speaker information included in the first speech content.
[0010] The present disclosure also provides a speech recognition system including a selection unit that selects first registration information to be used for speaker authentication from a database in which a plurality of registration information related to a registered speaker is registered; an acquisition unit that acquires a first speech voice signal obtained by having the speaker speak regarding the first registration information; a voice recognition unit that performs voice recognition on the first speech voice signal to acquire a first speech content of the speaker; and an authentication unit that determines whether or not speaker information included in the first speech content matches the first registration information; If it is determined that there is no match, second registration information different from the first registration information is selected, the acquisition unit acquires a second speech audio signal in which the speaker speaks regarding the second registration information, the speech recognition unit recognizes the second speech audio signal and acquires the second speech content of the speaker, and if it is determined that the speaker information included in the second speech content matches the second registration information, the authentication unit corrects the first registration information registered in the database to the speaker information included in the first speech content. [Effects of the Invention]
[0011] According to the present disclosure, it is possible to support management of user information used in voice authentication and suppress a decrease in the accuracy of voice authentication. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a diagram showing an example of a first use case of a voice authentication system according to an embodiment. [Figure 2] FIG. 10 is a diagram showing an example of a second use case of the voice authentication system according to the embodiment. [Figure 3] Block diagram showing an example of the internal configuration of an authentication analysis device [Figure 4] 10 is a flowchart illustrating an example of a registration procedure for an authentication analysis device according to an embodiment. [Figure 5] 1 is a flowchart illustrating an example of an authentication procedure of an authentication analysis device according to an embodiment. [Figure 6] A diagram explaining an example of updating registration information DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, with reference to the drawings as appropriate, detailed descriptions of embodiments specifically disclosing the registration information update method and registration information update device according to the present disclosure will be provided. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure and are not intended to limit the subject matter recited in the claims.
[0014] First, a first use case of the voice authentication system 100 according to the embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram illustrating an example of the first use case of the voice authentication system 100 according to the embodiment.
[0015] The voice authentication system 100 in the first use case is used, for example, when authenticating a speaker (user US) during automated voice guidance before receiving support from an operator (person) at a call center.
[0016] The voice authentication system 100 starts automatic voice guidance before establishing a communicable connection between the user-side call terminal UP1 and the operator-side call terminal OP1, and acquires a signal of the speech spoken by the user US during the automatic voice guidance (hereinafter referred to as "speech voice signal"). The voice authentication system 100 extracts features indicating the individuality of the user US from the acquired speech voice signal, and performs a matching process to match the extracted features with the features of the registered voice registered (stored) in the registered speaker database DB, and a matching process to match the speech content of the speech voice signal with the registered content registered (stored) in the registered speaker database DB, thereby performing user authentication.
[0017] The voice authentication system 100 in the first use case is configured to include at least a user-side communication terminal UP1, a network NW, an operator-side communication terminal OP1, an authentication analysis device P1, a registered speaker database DB, and an information display unit DP. The authentication analysis device P1 and the registered speaker database DB may be configured integrally. Similarly, the authentication analysis device P1 and the information display unit DP may be configured integrally. It goes without saying that the overall configuration of the voice authentication system 100 is not limited to the example shown in FIG. 1.
[0018] The user-side communication terminal UP1 is connected to the operator-side communication terminal OP1 and the authentication analysis device P1 via a network NW so as to be able to communicate wirelessly. Note that the wireless communication here refers to communication via a wireless LAN (Local Area Network) such as Wi-Fi (registered trademark).
[0019] The user-side communication terminal UP1 is realized by, for example, a notebook PC, a tablet terminal, a smartphone, or a telephone. The user-side communication terminal UP1 is a sound collection device equipped with a microphone (not shown), which collects the speech of the user US, converts it into a speech voice signal, and transmits the converted speech voice signal to the authentication analysis device P1 or the operator-side communication terminal OP1 via the network NW. The user-side communication terminal UP1 also acquires the audio signal of the automated voice guidance GD transmitted from the authentication analysis device P1 and outputs it from a speaker (not shown).
[0020] The network NW is an IP network or a telephone network that connects the user communication terminal UP1 and the operator communication terminal OP1 so that speech signals can be transmitted and received. Data is transmitted and received via wired or wireless communication.
[0021] The operator side communication terminal OP1 is connected to the user side communication terminal UP1 and the authentication analysis device P1 via wired or wireless communication so as to be able to transmit and receive data therebetween, and transmits and receives voice signals therebetween.
[0022] The authentication analysis device P1 is connected to the operator-side communication terminal OP1, the registered speaker database DB, and the information display unit DP so as to be able to transmit and receive data therebetween. Note that the authentication analysis device P1 may also be connected to the operator-side communication terminal OP1, the registered speaker database DB, and the information display unit DP so as to be able to communicate with each other via a network (not shown) in a wired or wireless manner.
[0023] The authentication analysis device P1 is realized by, for example, a notebook PC, a tablet terminal, or a smartphone. After being communicatively connected to the user-side communication terminal UP1, the authentication analysis device P1 executes automated voice guidance and user authentication. The authentication analysis device P1 acquires the speech signal of the user US transmitted from the user-side communication terminal UP1 via the network NW by automated voice guidance. When the authentication analysis device P1 authenticates the user US through user authentication, it switches the connection destination of the user-side communication terminal UP1 to the operator-side communication terminal OP1.
[0024] During voice registration, the authentication analysis device P1 associates the user US with identifiable user information (for example, the telephone number of the user's communication terminal UP1, the user's name, or user identification (ID)) and registers (stores) it in the registered speaker database DB.
[0025] Furthermore, during user authentication, the authentication analysis device P1 extracts speech features and analyzes the speech content based on the speech audio signal of the user US. The authentication analysis device P1 performs user authentication based on the speech features and user authentication based on the speech content. If user authentication based on the speech features and the speech content is successful, or if user authentication based on at least the speech features is unsuccessful, the authentication analysis device P1 generates an authentication result screen SC1 indicating the user authentication result, and transmits it to the information display unit DP for output. It goes without saying that the authentication result screen SC1 shown in FIG. 1 is merely an example and is not limited to this. The authentication result screen SC1 shown in FIG. 1 includes the user authentication result message "The voice matches that of XXXXX."
[0026] On the other hand, if the user authentication based on the utterance features is successful but the user authentication based on the utterance content is unsuccessful, the authentication analysis device P1 executes a process of updating the authentication words registered (stored) in the registered speaker database DB. The process of updating the authentication words will be described later.
[0027] The registered speaker database DB is a so-called storage, and is configured using a storage medium such as a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The registered speaker database DB stores (registers) user information of multiple users in association with speech features. The registered speaker database DB may be configured integrally with the authentication analysis device P1.
[0028] Here, user information is information about a user, and includes unchanging user information (hereinafter referred to as "unchanging information") and changeable user information (hereinafter referred to as "variable information"). Unchanging information in the present disclosure is user information that does not change after registration, such as the user's date of birth or membership number.
[0029] Variable information in the present disclosure is user information that can change after registration, such as a user name, address, or telephone number. User information classified as either unchanging information or variable information may be changed as appropriate depending on the usage scenario of the voice authentication system 100. In the present disclosure, an example will be described in which the user's date of birth and membership number are unchanging information, and other user information is variable information.
[0030] The information display unit DP is configured using, for example, a Liquid Crystal Display (LCD) or an organic electroluminescence (EL) display, and displays the authentication result screens SC1 and SC2 (see FIG. 2) transmitted from the authentication analysis device P1.
[0031] Next, a second use case of the voice authentication system 100A according to the embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the second use case of the voice authentication system 100A according to the embodiment.
[0032] As an example, the voice authentication system 100A in the second use case is used when a terminal device OP2 capable of picking up the speech of a user US requesting user authentication is placed in front of the user US, and the speaker (user US) is authenticated based on the speech picked up by the terminal device OP2.
[0033] The voice authentication system 100A acquires a speech signal of a speech uttered by a user US using a terminal device OP2. The voice authentication system 100A extracts features indicating the individuality of the user US from the acquired speech signal, and performs a matching process of matching the extracted features with features of registered speech registered (stored) in a registered speaker database DB, and a matching process of matching the speech content of the speech signal with the registered content registered (stored) in the registered speaker database DB, thereby performing user authentication.
[0034] The voice authentication system 100A in the second use case is configured to include at least a terminal device OP2, a network NWA, an authentication analysis device P1, and a registered speaker database DB. The authentication analysis device P1 and the registered speaker database DB may be configured integrally. It goes without saying that the overall configuration of the voice authentication system 100A is not limited to the example shown in FIG. 2.
[0035] The terminal device OP2 realizes the same functions as the operator-side call terminal OP1 and the information display unit DP in the first use case. The terminal device OP2 is connected to the authentication analysis device P1 via the network NWA so as to be able to send and receive data via wired or wireless communication, and executes data transmission and reception.
[0036] The terminal device OP2 is realized by, for example, a PC, a notebook PC, a tablet terminal, or a smartphone. The terminal device OP2 includes a display for displaying questions to prompt the user US to speak an authentication word and for displaying an authentication result screen SC2, and a microphone MK for collecting the voice of the user US. The microphone MK may be external to the terminal device OP2, or may be built into the terminal device OP2 and configured as an integrated unit. It goes without saying that the authentication result screen SC2 shown in FIG. 2 is merely an example and is not limited to this. The authentication result screen SC2 shown in FIG. 2 includes a message as the user authentication result, "You are Mr. / Ms. XX. Your identity has been verified."
[0037] The network NWA is connected between the user-side call terminal UP1 and the terminal device OP2 so as to enable wired or wireless communication between them, and transmits and receives data. The wireless network here corresponds to at least one of a wireless LAN such as Wi-Fi (registered trademark), a wireless WAN, and a mobile communication network such as 4G or 5G, and may also be other network configurations capable of wireless communication.
[0038] Next, an example of the internal configuration of the authentication analysis device P1 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the internal configuration of the authentication analysis device P1 according to an embodiment. The authentication analysis device P1 is configured to include at least a communication unit 20, a processor 21, and a memory 22.
[0039] The communication unit 20 is connected so as to enable data communication between the operator-side call terminal OP1 or the terminal device OP2 and the registered speaker database DB. The communication unit 20 acquires the speech signal of the user US from the operator-side call terminal OP1 that acquires the speech signal of the user US or the terminal device OP2 that collects the speech signal of the user US, and outputs the acquired speech signal to the processor 21.
[0040] The processor 21 is configured using a semiconductor chip on which at least one of electronic devices such as a central processing unit (CPU), a digital signal processor (DSP), a graphical processing unit (GPU), a field programmable gate array (FPGA), etc. The processor 21 functions as a controller that manages the overall operation of the authentication analysis device P1, and performs control processing for managing the operation of each part of the authentication analysis device P1, input / output processing of data between each part of the authentication analysis device P1, data arithmetic processing, and data storage processing.
[0041] Processor 21 uses programs and data stored in Read Only Memory (ROM) 221 of memory 22 to realize the functions of speech segment detection unit 211, question selection unit 212, speech recognition unit 213, speech recognition result analysis unit 214, feature extraction unit 215, similarity calculation unit 216, and registration information correction unit 217. Processor 21 uses Random Access Memory (RAM) 222 of memory 22 during operation, and temporarily stores data or information generated or acquired by processor 21 and each unit in RAM 222 of memory 22.
[0042] The processor 21 performs user authentication based on features indicating the individuality of the user US and user authentication based on the speech content of the user US, based on the speech signal of the user US. The processor 21 also modifies (updates) variable information among the user information registered in the registered speaker database DB.
[0043] The speech interval detection unit 211 analyzes the acquired speech audio signal and detects an utterance interval in which the user US is speaking. The speech interval detection unit 211 outputs an utterance audio signal corresponding to each utterance interval detected from the speech audio signal to the speech recognition unit 213 and the feature extraction unit 215. The speech interval detection unit 211 may also temporarily store the utterance audio signal of each utterance interval in the RAM 22B of the memory 22.
[0044] When registering the voice features and user information of the user US, the question selection unit 212 selects a question to prompt the user US to speak the user information to be registered in the registered speaker database DB. The question selection unit 212 outputs the selected question to the communication unit 20 and transmits it to the terminal device OP2, where it is displayed or output as voice.
[0045] Furthermore, during user authentication, the question selection unit 212 selects a question for prompting the user to speak any one of the multiple pieces of user information registered in the registered speaker database DB. The question selection unit 212 outputs the selected question to the communication unit 20 and transmits it to the terminal device OP2, where it is displayed or output as voice.
[0046] The speech recognition unit 213 uses machine learning, deep learning, or the like to analyze the speech audio signal corresponding to each speech section output from the speech section detection unit 211, and converts the speech content of the speech audio signal into text information (character information). The speech recognition unit 213 outputs the text information indicating the speech content to the speech recognition result analysis unit 214.
[0047] The speech recognition result analysis unit 214 compares the text information output from the speech recognition unit 213 with the user information registered in the registered speaker database DB and corresponding to the question selected by the question selection unit 212, and determines whether the content of the utterance matches the user information registered in the registered speaker database DB.
[0048] The feature extraction unit 215 extracts features indicating the individual characteristics of the user US from the speech sound signal output from the speech period detection unit 211. The feature extraction unit 215 outputs data of the extracted features to the similarity calculation unit 216.
[0049] The similarity calculation unit 216 compares the feature data output from the feature extraction unit 215 with the feature data of each of the utterances of multiple users registered in the registered speaker database DB, and calculates the similarity between the acquired feature data and the feature data of any of the users registered in the registered speaker database DB. Based on the calculated similarity, the similarity calculation unit 216 identifies the user corresponding to the speech voice signal and performs user authentication.
[0050] If the similarity calculation unit 216 determines that the user has been identified as a result of the user authentication, it generates an authentication result screen SC1 (see FIG. 1) including information about the identified user (i.e., the authentication result) and outputs it to the information display unit DP via the display I / F (Interface) 23. Furthermore, if the similarity calculation unit 216 determines that the user has been identified as a result of the user authentication, it generates an authentication result screen SC2 (see FIG. 2) including information about the identified user (i.e., the authentication result) and outputs it to the terminal device OP2 via the communication unit 20.
[0051] If the user authentication result is successful and the user information corresponding to the content of the user US's speech at the time of authentication does not match the user information (variable information) registered in the registered speaker database DB, the registered information correction unit 217 corrects (updates) the user information (variable information) registered in the registered speaker database DB to the user information corresponding to the content of the speech.
[0052] The memory 22 includes at least a ROM 221 that stores programs that define the various processes performed by the processor 21 and data used during execution of the programs, and a RAM 222 that serves as a work memory used when the various processes performed by the processor 21 are executed. The ROM 221 stores programs that define the various processes performed by the processor 21 and data used during execution of the programs. The RAM 222 temporarily stores data or information generated or acquired by the processor 21 (for example, speech audio signals before concatenation, speech audio signals after concatenation, speech features corresponding to each speech section before or after concatenation, etc.).
[0053] The display I / F 23 connects the processor 21 and the information display unit DP so that data communication is possible between them, and outputs to the information display unit DP an authentication result screen SC1 generated by the similarity calculation unit 216 of the processor 21. Note that the display I / F 23 is not an essential component and may be omitted.
[0054] Next, an example of a registration procedure for user information and features of a user's speech to be registered in the registered speaker database DB will be described with reference to Fig. 4 and Fig. 6. Fig. 4 is a flowchart illustrating an example of a registration procedure of the authentication analysis device P1 in the embodiment. Fig. 6 is a diagram illustrating an example of updating registered information.
[0055] 4 shows an example in which each of the processes in steps St11 to St13 is executed once, but is not limited to this. As will be described later, any of the processes in steps St11 to St13 may be executed multiple times until at least one invariant information and at least one variable information is acquired as user information to be registered in the registered speaker database DB.
[0056] The question selection unit 212 selects at least one question that is a registration word to be registered in the registered speaker database DB and that prompts the user US to speak user information. The question selection unit 212 transmits the selected question to the information display unit DP or the terminal device OP2, and causes it to be displayed or output as voice (St11).
[0057] The question selection unit 212 may select one or more questions. When multiple questions are selected, the voice authentication system 100, 100A may repeat the processes of steps St11 to St12 multiple times to obtain (utter) answers to each of the selected questions and acquire and register user information based on the uttered voice. The questions selected by the question selection unit 212 are one or more questions that prompt the user to utter (answer) user information to be registered in the registered speaker database DB, that is, at least one invariable information (e.g., the user's date of birth or membership number) and at least one variable information (e.g., the user's name, address, telephone number, etc.).
[0058] A specific example will be given of one or more questions for prompting the user US to utter (answer) at least one invariant piece of information and at least one variable piece of information. In the first processing (first step St11), the question selection unit 212 selects the question "What's your name?", which asks for the user's name, which is variable information; in the second processing (second step St11), it selects the question "What's your address?", which asks for the user's address, which is variable information; and in the third processing (third step St11), it selects the question "What's your date of birth?", which asks for the user's date of birth, which is variable information. As a result, in the processing of step St12 described below, the speech recognition result analysis unit 214 analyzes the content of the user US's utterance in response to each question, and can acquire the name and address of the user US, which are variable information, and the date of birth of the user US, which is invariant information.
[0059] The feature extraction unit 215 acquires a voice signal (hereinafter referred to as "registered utterance data") that captures the voice spoken by the user US, and extracts features that indicate the individuality of the user US from this registered utterance data (St12). Furthermore, the voice recognition unit 213 recognizes the content of the user US's utterance from the registered utterance data. The voice recognition result analysis unit 214 analyzes the recognized content of the utterance and extracts user information about the user US. Here, the extracted features may be features of a speech section corresponding to invariant information or variable information, or may simply be features that can be extracted from the entire speech section in which the user US is speaking.
[0060] The registered information correcting unit 217 associates the features extracted by the feature extracting unit 215 with the user information extracted by the speech recognition result analyzing unit 214, and registers (stores) them in the registered speaker database DB (St13). Note that, when the extracted features are features extracted from a speech section in which each user information (i.e., invariant information or variable information) is spoken, the registered information correcting unit 217 may associate the feature data with the user information from which the feature was extracted, and register them.
[0061] For example, the registered speaker database DB11 shown in Fig. 6 has registered therein, as user information, the name "XXX", the address "XX city, XX prefecture", the telephone number "000-1111-2222", the date of birth "2000 / 1 / 1", and the membership number "000111222". Here, the name, address, and telephone number information are variable information. The date of birth and membership number information are unchanging information. The registered information correction unit 217 registers (stores) each piece of user information in the registered speaker database DB11 based on the question selected by the question selection unit 212 and the analysis result of the speech content by the speech recognition result analysis unit 214.
[0062] As a result, the voice authentication system 100 according to the embodiment can acquire at least one invariant piece of information and at least one variable piece of information, and register these pieces of information as user information in association with the features of the speech of the user US.
[0063] Next, an example of a user authentication procedure based on the speech of the user US will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a flowchart illustrating an example of an authentication procedure of the authentication analysis device P1 in the embodiment.
[0064] The question selection unit 212 determines any variable information among the registration words registered in the registered speaker database DB as a first authentication word to be used for authentication. The question selection unit 212 selects at least one question for prompting the user US to utter the determined first authentication word. The question selection unit 212 transmits the selected question to the information display unit DP or the terminal device OP2, and causes it to be displayed or output as voice (St21).
[0065] The feature extraction unit 215 acquires a speech signal (hereinafter referred to as "first speech data") that captures the speech spoken by the user US as an answer to the question, and extracts features that indicate the individuality of the user US from this first speech data (St22).
[0066] The similarity calculation unit 216 reads the features of multiple users registered in the registered speaker database DB (St23). The similarity calculation unit 216 estimates which user US, who is the speaker, is registered in the registered speaker database DB based on the similarity between the feature of the user US extracted from the first utterance data and each of the features of the multiple users registered in the registered speaker database DB (St24).
[0067] The similarity calculation unit 216 performs a comparison to determine whether the feature amount of the user US extracted from the first utterance data matches the voice (feature amount) of a user registered in the registered speaker database DB (St25).
[0068] If the similarity calculation unit 216 determines in step St25 that the features of the user US match the voice (features) of a user registered in the registered speaker database DB (St25, YES), it identifies the user US as the user with whom these features match and requests the speech recognition unit 213 to recognize the content of the utterance of the first utterance data. The speech recognition result analysis unit 214 analyzes the recognized content of the utterance (St26).
[0069] On the other hand, if the similarity calculation unit 216 determines in step St25 that the features of the user US match the voice (features) of the user registered in the registered speaker database DB (St25, NO), it determines that the user US cannot be authenticated, generates an authentication failure screen (not shown) notifying the user of the failure of user authentication, transmits it to the information display unit DP or the terminal device OP2, and displays or outputs it as audio (St33).
[0070] The voice recognition result analysis unit 214 determines whether the analyzed utterance content is correct, that is, whether it is variable information indicated by the first authentication word (St27).
[0071] If the voice recognition result analysis unit 214 determines in step St27 that the analyzed utterance content is correct, that is, that it is the variable information indicated by the first authentication word (St27, YES), it determines that the user US has been authenticated, generates an authentication success screen (not shown) notifying the user of the successful user authentication, and sends it to the information display unit DP or the terminal device OP2 for display or audible output (St28).
[0072] On the other hand, if the voice recognition result analysis unit 214 determines in step St27 that the analyzed utterance content is incorrect, that is, that it is not the variable information indicated by the first authentication word (St27, NO), it requests the question selection unit 212 to select a second question.
[0073] Furthermore, if the similarity calculation unit 216 determines that the analyzed utterance content is incorrect (St27, NO), the similarity calculation unit 216 may set a threshold value for determining whether the user identified by the first utterance data and the user identified by the second utterance data are the same in a second matching process (step St31) described later, to be larger than the threshold value for determining whether the user identified by the first utterance data in the first matching process (step St25) is a user registered in the registered speaker database DB, in accordance with the security level of the installation where this user authentication is implemented. This allows the processor 21 to more effectively suppress erroneous user authentication and erroneous correction of user information.
[0074] The question selection unit 212 determines any invariant information or variable information among the registration words registered in the registered speaker database DB as a second authentication word to be used for authentication. The question selection unit 212 selects at least one question for prompting the user US to utter the determined second authentication word. The question selection unit 212 transmits the selected question to the information display unit DP or the terminal device OP2, and causes it to be displayed or output as voice (St29).
[0075] Note that, since other variable information may be changed in the same way, it is desirable that the second authentication word be unchanging information rather than variable information, but this is not limited to this. Furthermore, when determining the second authentication word as variable information, the variable information to be determined as the second authentication word may be selected based on the variable information determined as the first authentication word. For example, when the variable information determined as the first authentication word is a name, the question selection unit 212 may preferentially determine, as the second authentication word, variable information (e.g., a telephone number) that is unlikely to change along with the variable information determined as the first authentication word (here, a name).
[0076] The feature extraction unit 215 acquires a speech signal (hereinafter referred to as "second speech data") that captures the speech uttered by the user US as an answer to the question, and extracts features that indicate the individuality of the user US from this second speech data (St30).
[0077] The similarity calculation unit 216 determines whether the features of the user US extracted from the second utterance data match the voice (features) of the user identified in step St25. The voice recognition unit 213 recognizes the content of the utterance of the user US from the second utterance data (St31). The voice recognition result analysis unit 214 analyzes the recognized content of the utterance and determines whether the analyzed content of the utterance is correct, that is, whether it is invariant information or variable information indicated by the second authentication word (St31).
[0078] In step St31, if the similarity calculation unit 216 determines that the features of the user US extracted from the second speech data by the similarity calculation unit 216 match the voice (features) of the identified user and if the speech content analyzed by the voice recognition result analysis unit 214 is determined to be correct (St31, YES), the similarity calculation unit 216 requests the registration information correction unit 217 to correct (update) the user information of this user US.
[0079] The registration information correction unit 217 corrects the user information (variable information) registered in the registered speaker database DB, which corresponds to the first authentication word, to the utterance content analyzed from the first utterance data (St32). Furthermore, the registration information correction unit 217 adds the feature amount extracted from the first utterance data to the feature amount of the user US registered in the registered speaker database DB, and averages and stores (registers) the result. This allows the authentication analysis device P1 to increase the feature amount of the user US, thereby more efficiently improving the user authentication accuracy.
[0080] For example, if the user information corresponding to the first authentication word among the user information registered in the registered speaker database DB11 is an address "XX prefecture, △△ city," and the utterance content analyzed from the first utterance data is "XX prefecture, □□ city," the registered information correcting unit 217 corrects (updates) the address of the user US registered in the registered speaker database DB11 from "XX prefecture, △△ city" to "XX prefecture, □□ city." The registered speaker database DB12 shows a state in which the address information, which is variable information among the user information shown in the registered speaker database DB11, has been corrected (updated).
[0081] As described above, the voice authentication system 100 according to the embodiment can acquire at least one invariant information and at least one variable information, and register these pieces of information as user information in association with features of the speech of the user US. As a result, even when performing user authentication using user information that may change in the future due to, for example, moving, changing a telephone number, or marriage, the voice authentication system 100 can effectively suppress authentication failures based on the changed user information (variable information) and can automatically update the user information (variable information). Therefore, the user can continue to receive user authentication without having to update the user information in advance.
[0082] Furthermore, when the voice authentication system 100 performs authentication using user information, which is unchanging information, in the second authentication (processing of steps St29 to St31), it can improve the accuracy of identifying the user, thereby more effectively suppressing authentication failures based on changed user information (variable information) and more accurately performing automatic updates of the user information (variable information).
[0083] In the above description, the process of registering at least one invariant information and at least one variable information in the voice authentication system 100 according to the embodiment does not have to be limited to a registration method using voice recognition. For example, the voice authentication system 100 may accept input and registration of at least one invariant information and at least one variable information by an input operation by the user US to the user-side call terminal UP1 or the terminal device OP2.
[0084] Furthermore, the voice authentication system 100 according to the embodiment has been shown as an example of performing 1:N (N: an integer of 2 or more) authentication to determine whether a feature of one user US is one of a plurality of users, but it may also perform 1:1 authentication to determine whether the user US is a predetermined user. When performing 1:1 authentication, the voice authentication system 100 may determine, in the first authentication (processing of step St25), whether the user US matches the phone number, Internet Protocol (IP) address, first authentication password, or user identified by reading a membership card or the like of the user side call terminal UP1.
[0085] (Addendum) The above description of each embodiment discloses the following techniques.
[0086] (Technology 1) A registration information update method performed by at least one computer (authentication analysis device P1), Selecting first registration information (user information that is a first authentication word and is variable information) to be used for authenticating a speaker (user) from a database (registered speaker database DB11) in which a plurality of registration information related to a registered speaker (user) is registered; A first speech signal is acquired by having the speaker (user) speak about the first registration information (user information that is variable information); performing speech recognition on the first speech voice signal to obtain the first speech content of the speaker (user); If it is determined that the speaker (user) information included in the first utterance content does not match the first registered information, second registered information (user information that is variable information, being a second authentication word) different from the first registered information (user information that is variable information) is selected, and a second speech voice signal is acquired by having the speaker (user) speak the second registered information; performing speech recognition on the second speech audio signal to obtain second speech content of the speaker (user); When it is determined that the speaker (user) information included in the second utterance content matches the second registered information, the first registered information (user information which is variable information) registered in the database (registered speaker database DB11) is corrected to the speaker (user) information included in the first utterance content. How to update your registration information. As a result, even when the computer (authentication analysis device P1) performs user authentication using user information that may be changed in the future, it can effectively prevent authentication failures based on changed user information (variable information) and automatically update old user information (variable information) to the latest user information. Therefore, the computer (authentication analysis device P1) can more effectively prevent a decrease in voice authentication accuracy caused by unmodified (unupdated) user information used for user authentication. Users can continue to receive user authentication without having to update their user information in advance.
[0087] (Technology 2) The database (registered speaker database DB11) registers the plurality of registration information of the registered speaker (user) in association with the feature amount of the voice of the registered speaker (user), extracting a first feature of the speaker (user) based on the first speech signal; If it is determined that the first feature matches with a feature of any registered speaker (user) registered in the database (registered speaker database DB11), the speaker (user) is made to speak the second registration information. (Technical 1) describes a method for updating registration information. This allows the computer (authentication analysis device P1) to perform authentication using the second authentication word only if the voice is that of one of the pre-registered users, thereby more effectively preventing erroneous corrections to user information.
[0088] (Technology 3) extracting a second feature of the speaker (user) based on the second speech signal when it is determined that the first feature matches the feature of the registered speaker (user); When it is determined that the second feature matches the feature of the registered speaker (user), the first registered information (user information which is variable information) is corrected to the speaker (user) information. (Technical 2) describes a method for updating registration information. This allows the computer (authentication analysis device P1) to more effectively prevent erroneous correction of user information by correcting the user information corresponding to the first authentication word only when the user matched in authentication using the first speech data matches the user matched in authentication using the second speech data.
[0089] (Technology 4) The first enrollment information is variable information that can be changed by the enrollment speaker (user) after being registered in the database (enrollment speaker database DB11). A registration information updating method according to any one of (Technology 1) to (Technology 3). As a result, the computer (authentication analysis device P1) can modify the user information corresponding to the first authentication word if the user authentication result based on the second speech data is successful, even if the voice recognition result based on the first authentication word is unsuccessful in authentication based on the first speech data, by changing the user information.
[0090] (Technology 5) The second enrollment information is invariant information that is not changed by the enrollment speaker (user) after being registered in the database (enrollment speaker database DB11). A registration information update method according to any one of (Technology 1) to (Technology 4). As a result, the computer (authentication analysis device P1) can modify the user information corresponding to the first authentication word if the user authentication result based on the second speech data is successful, even if the voice recognition result based on the first authentication word is unsuccessful in authentication based on the first speech data, by changing the user information.
[0091] (Technology 6) a selection unit (question selection unit 212) that selects first registration information (user information that is a first authentication word and is variable information) to be used for speaker (user) authentication from a database (registered speaker database DB11) in which a plurality of registration information related to registered speakers (users) is registered; an acquisition unit (communication unit 20) that acquires a first speech voice signal obtained by having the speaker (user) speak about the first registration information (user information that is variable information); a voice recognition unit 213 that performs voice recognition on the first utterance voice signal to acquire the first utterance content of the speaker (user); an authentication unit (similarity calculation unit 216) that determines whether speaker (user) information included in the first utterance content matches the first registered information; When the authentication unit determines that the speaker (user) information included in the first utterance content does not match the first registered information, the selection unit (question selection unit 212) selects second registered information (user information that is a second authentication word and is variable information or invariant information) that is different from the first registered information, The acquisition unit (communication unit 20) acquires a second speech voice signal obtained by having the speaker (user) speak regarding the second registration information, the speech recognition unit 213 performs speech recognition on the second speech signal to acquire the second speech content of the speaker (user); When it is determined that the speaker (user) information included in the second utterance content matches the second registered information, the authentication unit (similarity calculation unit 216) modifies the first registered information (user information that is variable information) registered in the database (registered speaker (user) database DB11) to the speaker (user) information included in the first utterance content. Registration information update device (authentication analysis device P1). As a result, even when performing user authentication using user information that may be changed in the future, the registration information update device (authentication analysis device P1) can effectively prevent authentication failures based on changed user information (variable information) and automatically update old user information (variable information) to the latest user information. Therefore, the computer (authentication analysis device P1) can more effectively prevent a decrease in voice authentication accuracy caused by unmodified (unupdated) user information used for user authentication. Users can continue to receive user authentication without having to update their user information in advance.
[0092] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention. [Industrial Applicability]
[0093] The present disclosure is useful for a voice registration device and a voice registration method that support management of user information used in voice authentication and suppress a decrease in voice authentication accuracy. [Explanation of symbols]
[0094] 20 Communications Department 21 processors 22 Memory 100,100A Voice Authentication System 211 Speech Activity Detection Unit 212 Question Selection Section 213 Voice Recognition Unit 214 Speech recognition result analysis unit 215 Feature Extraction Unit 216 Similarity calculation part 217 Registration Information Correction Section DB, DB11, DB12 registered speaker database DP information display section MK Microphone OP Operator OP1 Operator side call terminal OP2 terminal device P1 Authentication Analysis Device SC1, SC2 authentication result screen UP1 User side call terminal US users
Claims
1. 1. A method for updating registration information performed by at least one computer, comprising: selecting a first enrollment information item to be used for speaker authentication from a database in which a plurality of enrollment information items relating to enrollment speakers are registered; acquiring a first speech audio signal obtained by having the speaker speak regarding the first enrollment information; performing speech recognition on the first speech audio signal to obtain a first speech content of the speaker; If it is determined that the speaker information included in the first utterance content does not match the first enrollment information, second enrollment information different from the first enrollment information is selected, and a second speech audio signal is acquired by having the speaker speak regarding the second enrollment information; performing speech recognition on the second speech audio signal to obtain second speech content of the speaker; when it is determined that the speaker information included in the second utterance content matches the second registered information, correcting the first registered information registered in the database to the speaker information included in the first utterance content; How to update your registration information.
2. the database registers the plurality of pieces of registration information of the registration speakers in association with the feature quantities of the speech of the registration speakers; extracting a first feature of the speaker based on the first speech signal; when it is determined that the first feature matches with a feature of any registered speaker registered in the database, the speaker is made to speak the second registered information; The registration information updating method according to claim 1 .
3. extracting a second feature of the speaker based on the second speech signal when it is determined that the first feature matches the feature of the registered speaker; when it is determined that the second feature matches the feature of the registered speaker, correcting the first registered information to the speaker information; The registration information updating method according to claim 2.
4. the first enrollment information is variable information that can be changed by the enrollment speaker after being registered in the database; The registration information updating method according to claim 1 .
5. the second enrollment information is invariant information that is not changed by the enrollment speaker after being registered in the database; The registration information updating method according to claim 1 .
6. a selection unit that selects first enrollment information to be used for speaker authentication from a database in which a plurality of enrollment information related to enrolled speakers is registered; an acquisition unit that acquires a first speech sound signal obtained by having the speaker speak regarding the first registration information; a speech recognition unit that recognizes the first speech signal and acquires the first speech content of the speaker; an authentication unit that determines whether speaker information included in the first utterance content matches the first registration information; the selection unit, when the authentication unit determines that speaker information included in the first utterance content does not match the first enrollment information, selects second enrollment information different from the first enrollment information; the acquiring unit acquires a second speech sound signal obtained by making the speaker speak about the second enrollment information; the speech recognition unit performs speech recognition on the second speech voice signal to acquire second speech content of the speaker; When it is determined that the speaker information included in the second utterance content matches the second registered information, the authentication unit corrects the first registered information registered in the database to the speaker information included in the first utterance content. Registration information update device.
Citation Information
Patent Citations
Information processing device and program
JP2022007620A