Registration information updating method and registration information updating apparatus
The voice authentication method improves accuracy and security by employing two authentication words, one pre-registered and one randomly generated, to enhance user verification and thwart impersonation.
Patent Information
- Application Number
- JP2024073767
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-11-12
AI Technical Summary
Conventional voice authentication methods face challenges in maintaining high accuracy while preventing spoofing and impersonation, particularly when using recorded or synthesized speech.
A voice authentication method that involves two distinct authentication words, where the first word is pre-registered and the second word is randomly generated, combining speaker estimation processes to verify user identity.
Enhances voice authentication accuracy and effectively prevents spoofing by using randomly generated words that are difficult to predict, reducing impersonation attempts.
Smart Images

Figure 2025168908000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a registration information updating method and a registration information updating device. [Background technology]
[0002] Patent document 1 discloses an information processing device that acquires each of multiple voices uttered by a user in a series of multiple operations, performs a second authentication to confirm that a voice corresponding to a specific operation that is part of the multiple operations among the multiple voices is the voice of the user authenticated in the first authentication, and once the second authentication is performed, executes processing according to the multiple operations associated with the recognition results of the multiple voices. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-197060 Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure has been devised in view of the above-described conventional situation, and aims to provide a voice authentication method and a voice authentication device that more effectively prevent a decrease in voice authentication accuracy and impersonation by others in voice authentication. [Means for solving the problem]
[0005] The present disclosure provides a voice authentication method performed by at least one computer, the voice authentication method comprising: acquiring first speech data in which a speaker speaks a first authentication word; performing first speaker estimation to estimate which of the registered speakers the speaker is, based on a first feature of the speaker extracted from the first speech data and a feature of at least one registered speaker registered in advance; acquiring second speech data in which the speaker is spoken a second authentication word different from the first authentication word; and, if it is determined that the speech content based on the second speech data matches the second authentication word, performing second speaker estimation to estimate which of the registered speakers the speaker is, based on a second feature of the speaker extracted from the second speech data and the feature of the registered speaker; and authenticating the speaker based on whether the registered speaker estimated by the first speaker estimation matches the registered speaker estimated by the second speaker estimation.
[0006] The present disclosure also provides a speech recognition system including an output unit that outputs a first authentication word to be spoken by a speaker and a second authentication word different from the first authentication word; an acquisition unit that acquires first spoken voice data in which the speaker is made to speak the first authentication word and second spoken voice data in which the speaker is made to speak the second authentication word; and a first speech recognition system that estimates which of the registered speakers the speaker is a registered speaker of, based on a first feature of the speaker extracted from the first spoken voice data and a feature of at least one registered speaker that is registered in advance. Provided is a voice authentication device comprising: an estimation unit that performs second speaker estimation to estimate which registered speaker among the registered speakers the speaker is, based on second features of the speaker extracted from the second speech voice data and features of the registered speakers, when it is determined that the speaker estimation, the speech content based on the second speech voice data, and the second authentication word match; and an authentication unit that authenticates the speaker based on whether the registered speaker estimated by the first speaker estimation matches the registered speaker estimated by the second speaker estimation. [Effects of the Invention]
[0007] According to the present disclosure, in voice authentication, it is possible to more effectively prevent a decrease in voice authentication accuracy and spoofing by others. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of a first use case of a voice authentication system according to an embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example of a second use case of the voice authentication system according to the embodiment. [Figure 3] FIG. 1 is a block diagram showing an example of the internal configuration of an authentication device according to an embodiment. [Figure 4] 1 is a flowchart showing an example of a user registration procedure for an authentication device according to an embodiment. [Figure 5] 1 is a flowchart showing an example of a user authentication procedure of an authentication device according to an embodiment. [Figure 6] FIG. 10 is a diagram showing an example of a second authentication word in the embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] (Background to this disclosure) Conventional voice authentication has two methods: the first, in which features that indicate the user's individuality are extracted from the registration voice in which a registration word is spoken at the time of registration and registered in advance, and then the user is authenticated by having the user speak the same word as the registration word at the time of authentication; and the second, in which the user is authenticated by having the user speak a word different from the registration word.
[0010] However, the first authentication method requires the user to speak the same words as the registered words, which allows for highly accurate voice authentication, but has the problem of making it possible for spoofing to occur using a recording of the user's speech. The second authentication method requires the user to speak words different from the registered words, which makes it difficult to make spoofing to occur using a recording of the user's speech, but there is a possibility that authentication accuracy will decrease. Therefore, there has been a demand for a voice authentication method that provides highly accurate voice authentication while preventing spoofing by others.
[0011] Incidentally, in Patent Document 1, a voice operation of a user authenticated by login authentication (first authentication) is accepted, and based on the voice (voice indicating an authentication word) generated during this voice operation, additional authentication (second authentication) is performed to confirm whether the user who generated the voice operation is the user authenticated by login authentication (first authentication), thereby preventing impersonation by a third party. However, in the method described in Patent Document 1, the authentication word (word) is common to all users, so there is a possibility of impersonation by voice imitation, voice synthesis, recording, or the like.
[0012] Hereinafter, with reference to the accompanying drawings as appropriate, detailed descriptions of embodiments specifically disclosing a voice authentication method and voice authentication device according to the present disclosure will be provided. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter recited in the claims.
[0013] First, a first use case of the voice authentication system 100 according to the embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram illustrating an example of the first use case of the voice authentication system 100 according to the embodiment.
[0014] The voice authentication system 100 in the first use case is an example that is performed locally without using a network. As an example, the voice authentication system 100 shown in Fig. 1 performs speaker (user) authentication using an information display device DP1 and a sound collection device MC1 that can be used by one or more users, and an authentication device P1 that is communicably connected between the information display device DP1 and the sound collection device MC1.
[0015] The voice authentication system 100 presents (displays or outputs) speech words to be spoken by a user using an information display device DP1, and acquires speech data (speech signal) of the speech spoken by the user using a sound collection device MC1. The voice authentication system 100 analyzes the user's speech content from the speech voice data and performs a speech content authentication process to determine whether the analyzed speech content matches the speech words, and a user estimation process to extract features indicating the user's individuality from the acquired speech voice data and compare the extracted features with features of registered voices registered (stored) in a registered speaker database DB to estimate the user who is the speaker. The voice authentication system 100 performs the above-mentioned speech content authentication process and user estimation process as a single authentication process, and performs user authentication based on whether the user estimated in the first authentication process matches the user estimated in the second authentication process.
[0016] The voice authentication system 100 in the first use case includes an information display device DP1, a sound collection device MC1, an authentication device P1, and a registered speaker database DB. It goes without saying that the overall configuration of the voice authentication system 100 is not limited to the example shown in Fig. 1. For example, the information display device DP1 and the sound collection device MC1 may be configured integrally and may be realized by a personal computer (hereinafter referred to as "PC"), a notebook PC, a tablet terminal, or the like. Furthermore, the authentication device P1 and the registered speaker database DB may be configured integrally.
[0017] The information display device DP1 is configured using, for example, a Liquid Crystal Display (LCD) or an organic electroluminescence (EL) display. The information display device DP1 is connected to the authentication device P1 so as to be able to send and receive data via wired or wireless communication. The information display device DP1 displays the speech words that are sent from the authentication device P1 to be spoken by the user, or questions that prompt the user to speak the speech words, and displays an authentication screen (not shown) that notifies the user of the result of user authentication sent from the authentication device P1.
[0018] The sound collection device MC1 is a microphone for collecting the user's speech. The sound collection device MC1 may be attached externally to the information display device DP1, or may be built into the information display device DP1 and configured as an integral part of the information display device DP1.
[0019] The authentication device P1 is connected to the information display device DP1, the sound collection device MC1, and the registered speaker database DB so that data can be transmitted and received among them. The authentication device P1 is realized by a PC, a laptop computer, a tablet terminal, a server, a cloud server, or the like. The authentication device P1 performs user registration, which registers features extracted from the user's speech data in the registered speaker database DB, and user authentication, which authenticates the user based on a comparison with the features of at least one user registered in the registered speaker database DB.
[0020] The registered speaker database DB is a so-called storage, and is configured using a storage medium such as a flash memory, a hard disk drive (HDD), or a solid state drive (SSD). The registered speaker database DB stores (registers) user information of multiple users and speech features in association with each other. The registered speaker database DB may be configured integrally with the authentication device P1. The user information is information related to a user, such as the user's name, date of birth, affiliation information, or user identification (ID).
[0021] Next, a second use case of the voice authentication system 100A according to the embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the second use case of the voice authentication system 100A according to the embodiment.
[0022] As an example, the voice authentication system 100A in the second use case authenticates a speaker (user) using a user terminal U1 owned by a user who desires user authentication and an authentication device P1 communicatively connected to the user terminal U1 via a network NW.
[0023] The voice authentication system 100A presents (displays or outputs) speech words to be spoken by a user terminal U1 (information display device DP), and acquires speech data (speech signal) of the speech spoken by the user via the user terminal U1. The voice authentication system 100A analyzes the content of the user's speech from the speech data and performs a speech content authentication process to determine whether the analyzed content of the speech matches the speech words, and also performs a user estimation process to extract features indicating the user's individuality from the acquired speech data and compare the extracted features with features of registered speech registered (stored) in a registered speaker database DB to estimate the user who is the speaker. The voice authentication system 100A performs the speech content authentication process and the user estimation process as a single authentication process, and performs user authentication based on whether the user estimated in the first authentication process matches the user estimated in the second authentication process.
[0024] The user terminal U1 is connected to the authentication device P1 so as to be able to communicate wirelessly with the authentication device P1. The wireless communication here refers to communication via a wireless local area network (LAN) such as Wi-Fi (registered trademark).
[0025] The user terminal U1 is realized by, for example, a PC, notebook PC, tablet terminal, smartphone, or the like owned by the user. The user terminal U1 integrally includes a sound collection unit MC2 that realizes the same functions as the sound collection device MC1 in the first use case example, and a display unit DP2 that realizes the same functions as the information display device DP in the first use case example. Note that the functions of the sound collection unit MC2 or the display unit DP2 may be realized by a device externally attached to the user terminal U1.
[0026] In the second use case, the authentication device P1 is communicably connected to the user terminal U1 and the registered speaker database DB.
[0027] Next, an example of the internal configuration of the authentication device P1 will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the internal configuration of the authentication device P1 in the embodiment. In the following description, the information display device DP indicates the information display device DP1 in the voice authentication system 100, and indicates the display unit DP2 provided in the user terminal U1 in the voice authentication system 100A. The sound collection device MC indicates the sound collection device MC1 in the voice authentication system 100, and indicates the sound collection unit MC2 provided in the user terminal U1 in the voice authentication system 100A.
[0028] The authentication device P1 includes at least a communication unit 10, a processor 11, a memory 12, and a display I / F 13.
[0029] The communication unit 10 is connected to the information display device DP, the sound collection device MC, and the registered speaker database DB so as to be able to communicate data with each other. The communication unit 10 transmits various screens output from the processor 11 to the information display device DP, acquires user speech data transmitted from the sound collection device MC, and outputs it to the processor 11.
[0030] The processor 11 is configured using a semiconductor chip on which at least one of electronic devices such as a central processing unit (CPU), a digital signal processor (DSP), a graphical processing unit (GPU), a field programmable gate array (FPGA), etc. The processor 11 functions as a controller that manages the overall operation of the authentication device P1, and performs control processing for managing the operation of each part of the authentication device P1, data input / output processing between each part of the authentication device P1, data calculation processing, and data storage processing.
[0031] The processor 11 realizes the user registration function and the user authentication function by using the programs and data stored in the read only memory (ROM) 221 of the memory 12. The processor 11 uses the random access memory (RAM) 122 of the memory 12 during operation, and temporarily stores data or information generated or acquired by the processor 11 and each unit in the RAM 122 of the memory 12.
[0032] The memory 12 includes at least a ROM 121 that stores, for example, a program that defines the various processes performed by the processor 11 and data used during execution of the program, and a RAM 122 that serves as a work memory used when executing the various processes performed by the processor 11. The ROM 121 stores the program that defines the various processes performed by the processor 11 and the data used during execution of the program. The RAM 122 temporarily stores data or information generated or acquired by the processor 11.
[0033] The memory 12 stores various data used to generate second authentication words, which will be described later. The memory 12 includes a word dictionary 123 that stores word data. The word dictionary 123 stores data on a plurality of words. The memory 12 also stores fixed phrases and the like that are used to generate second authentication words, which will be described later.
[0034] The display I / F 13 connects the processor 11 and the information display device DP so that data communication is possible between them, and outputs to the information display device DP an authentication screen (not shown) showing the authentication result generated by the processor 11. Note that the display I / F 23 is not an essential component and may be omitted.
[0035] Next, an example of a procedure for registering user information and features of a user's speech to be registered in the registered speaker database DB will be described with reference to Fig. 4. Fig. 4 is a flowchart illustrating an example of a user registration procedure in the authentication device P1 according to the embodiment.
[0036] 4 shows an example in which each of the processes in steps St11 to St13 is executed once, but is not limited to this. The processes in steps St11 to St13 may be repeatedly executed a number of times corresponding to the number of registration words used as first authentication words during user authentication, which is the number of pieces of user information registered in the registered speaker database DB.
[0037] The processor 11 selects at least one registration word to prompt the user to speak user information (e.g., the user's name, date of birth, affiliation information, or user ID) to be registered in the registered speaker database DB. The registration word here refers to at least one piece of text data or question data. The processor 11 generates a screen including the selected text or question, and transmits it to the information display device DP for display (St11).
[0038] Here, one or more questions may be selected by the processor 11. When multiple texts or questions are selected, the voice authentication system 100, 100A may repeat the processes of steps St11 to St12 multiple times to obtain answers (utterances) to each of the multiple selected texts or questions, and to acquire and register user information based on the uttered voices.
[0039] The processor 11 acquires speech data (hereinafter referred to as "registration speech data") transmitted from the sound collection device MC and capturing the speech of the user uttering the registration words (St12). The processor 11 extracts and acquires features indicative of the user's individuality from this registration speech data (St12). Here, the extracted features may be features of a speech section in which the registration words are spoken, or may simply be features extractable from the entire speech section in which the user is speaking. The processor 11 also recognizes the content of the user's speech from the registration speech data, analyzes the recognized content of the speech, and extracts user information about the user.
[0040] The processor 11 associates the acquired feature with the user information and registers (stores) it in the registered speaker database DB (St14). Note that, when the extracted feature is extracted from an utterance section in which each registration word is spoken, the processor 11 may register the feature data and the user information from which the feature was extracted in association with each other.
[0041] As described above, the voice authentication systems 100 and 100A according to the embodiments can acquire and register user information, which is a registration word, and the feature of the user who speaks the registration word. This allows the voice authentication systems 100 and 100A to improve the accuracy of user authentication using the registration word.
[0042] Next, an example of a user authentication procedure based on a user's utterance will be described with reference to Fig. 5 and Fig. 6. Fig. 5 is a flowchart showing an example of a user authentication procedure of the authentication device P1 in the embodiment. Fig. 6 is a diagram showing an example of a second authentication word in the embodiment.
[0043] Processor 11 determines the registration word used in step St11 during user registration as the first authentication word. Processor 11 selects at least one piece of text data or question data for prompting the user to speak the determined first authentication word. Processor 11 generates a screen including the selected text or question, and transmits the generated screen to information display device DP for display (St21). Note that if there are multiple registration words, processor 11 may select any one of the registration words.
[0044] The processor 11 acquires speech data (hereinafter referred to as "first speech data") that is transmitted from the sound collection device MC and that is obtained by collecting the speech of the user uttering the first authentication word (St22). The processor 11 extracts and acquires features that indicate the individuality of the user from the first speech data (St22).
[0045] The processor 11 estimates which of the users registered in the registered speaker database DB the user to be authenticated is based on the features of multiple users registered in the registered speaker database DB and the feature of the user extracted from the first speech voice data (St23).The processor 11 executes a determination (first speaker estimation) as to whether the feature of the user matches or is similar to the feature of any one of the users registered in the registered speaker database DB (St24).
[0046] In step St24, if the processor 11 determines that the user's features match or are similar to the features of any one of the users registered in the registered speaker database DB (St24, YES), the processor 11 generates a second authentication word that is different from the registration word. The processor 11 selects at least one piece of text data or question data for prompting the user to speak the determined second authentication word. The processor 11 generates a screen including the selected text or question, and transmits the generated screen to the information display device DP for display (St25).
[0047] The second authentication word will now be described with reference to Fig. 6. The second authentication word described above is a word that is not registered (stored) in the registered speaker database DB, and is not a specific word determined in advance like user information, but is a word that is frequently changed for each authentication.
[0048] For example, if the date and time information for user authentication is "3 / 20 13:56", the second authentication word may be "March 20th 13:56", or it may be a word made up of a combination of multiple words randomly selected from the word dictionary 123, a standard phrase containing multiple words randomly selected from the word dictionary, a sentence containing information about the date and time or place for user authentication, or a word obtained by combining these pieces of information.
[0049] Here, the word formed by combining a plurality of words randomly selected from the word dictionary 123 is the word "dog, apple, chocolate" obtained by randomly selecting a plurality of words (for example, "dog," "apple," and "chocolate") from the word dictionary 123 and arranging these selected words. Furthermore, the fixed phrase including a plurality of words randomly set from the word dictionary 123 may be, for example, the sentence "Let's eat apples and chocolate today" obtained by randomly selecting "food 1: apples" and "food 2: chocolate" related to food from the word dictionary and inserting the selected "food 1: apples" and "food 2: chocolate" into <food 1> and <food 2>, respectively, of the fixed phrase "Let's eat <food 1> and <food 2> today." Furthermore, a template containing multiple words randomly selected from the word dictionary 123 may be a sentence that does not have the meaning of a template, for example, it may be the sentence "Tomorrow I will play with dog and chocolate" obtained by randomly selecting arbitrary "word 1: dog" and "word 3: chocolate" from the word dictionary 123 and inserting "word 1: dog" and "word 3: chocolate" into <word 1> and <word 3>, respectively, of the template "Tomorrow I will play with <word 1> and <word 3>."
[0050] Furthermore, the sentence relating to the date, time or location where user authentication is performed may be, for example, if user authentication is performed in the location "Tokyo" and the weather in the location "Tokyo" is "sunny" and the temperature is "20°C," a sentence using this location, weather and temperature information may be, "The weather in Tokyo is currently sunny and the temperature is 20°C."
[0051] On the other hand, if the processor 11 determines in step St24 that the user's features do not match or are similar to the features of any user registered in the registered speaker database DB (St24, NO), it generates an authentication screen (not shown) notifying that the result of user authentication is a failure (i.e., authentication failed), sends it to the information display device DP, and displays it (St26).
[0052] The processor 11 acquires speech data (hereinafter referred to as "second speech data") that is transmitted from the sound collection device MC and is obtained by collecting the speech of the user uttering the second authentication word (St27).
[0053] The processor 11 performs a voice authentication process on the acquired second speech data to recognize the user's speech content, analyzes the recognized speech content, and determines whether the analyzed speech content matches the second authentication word (St28).
[0054] If the processor 11 determines in step St28 that the analyzed speech content does not match the second authentication word (St28, NO), it generates an authentication screen (not shown) notifying that the result of user authentication is a failure (i.e., authentication failed), sends it to the information display device DP, and displays it (St26).
[0055] On the other hand, if the processor 11 determines in step St28 that the analyzed speech content matches the second authentication word (St28, YES), it extracts features indicating the individuality of the user from the second speech voice data and estimates which user registered in the registered speaker database DB the user to be authenticated is based on the features of multiple users registered in the registered speaker database DB and the user features extracted from the second speech voice data (St29).The processor 11 executes a determination (second speaker estimation) as to whether the user estimated using the first speech voice data matches the user estimated using the second speech voice data (St30).
[0056] If the processor 11 determines in step St30 that the user estimated using the first speech data matches the user estimated using the second speech data (St30, YES), it generates an authentication screen (not shown) notifying that the result of user authentication is successful (i.e., authentication successful), sends it to the information display device DP, and displays it (St31).
[0057] On the other hand, if the processor 11 determines in step St30 that the user estimated using the first speech voice data does not match the user estimated using the second speech voice data (St30, NO), it generates an authentication screen (not shown) notifying that the result of user authentication is a failure (i.e., authentication failed), sends it to the information display device DP, and displays it (St26).
[0058] As described above, the authentication device P1 in the embodiment can authenticate a user estimated using the first speech voice data and a user estimated using the second speech voice data. Furthermore, the authentication device P1 can perform voice authentication with higher accuracy by setting the first authentication word to the same word as a pre-registered registration word when acquiring the first speech voice data.
[0059] Furthermore, when acquiring the second spoken voice data, the authentication device P1 determines the second authentication word to be a combination of a plurality of randomly selected words or a second authentication word including a plurality of randomly selected words, or a second authentication word including information such as the location, date, time, or weather at the time of user authentication, thereby more effectively preventing impersonation by others using recorded data. Furthermore, by determining the second authentication word as described above when acquiring the second spoken voice data, the authentication device P1 can request the user to speak in a semi-free manner but including a larger number of phonemes, thereby improving authentication accuracy.
[0060] (Addendum) The above description of each embodiment discloses the following techniques.
[0061] (Technology 1) A voice authentication method performed by at least one computer (authentication device P1), acquiring first speech data obtained by having a speaker (i.e., a user) speak a first authentication word; performing a first speaker estimation to estimate which of the registered speakers the speaker is, based on a first feature of the speaker extracted from the first speech voice data and a feature of at least one registered speaker (i.e., a user registered in a registered speaker database DB) registered in advance; acquiring second speech data by having the speaker speak a second authentication word different from the first authentication word; if it is determined that the speech content based on the second speech voice data matches the second authentication word, a second speaker estimation is performed to estimate which of the registered speakers the speaker is, based on a second feature of the speaker extracted from the second speech voice data and the feature of the registered speaker; authenticating the speaker based on whether or not the enrollment speaker estimated by the first speaker estimation matches the enrollment speaker estimated by the second speaker estimation; Voice authentication method. With this configuration, the authentication device P1 determines whether the user estimated from each of the first speech voice data and the second speech voice data, in which the speaker is made to speak two different authentication words (a first authentication word and a second authentication word), matches, thereby more effectively suppressing impersonation by others and performing more accurate voice authentication.
[0062] (Technology 2) generating the second speech voice data when it is estimated by the first speaker estimation that the speaker is one of the registered speakers; A voice authentication method described in (Technology 1). With this configuration, the authentication device P1 generates the second authentication word only when the user is estimated in the first speaker estimation, and performs the second speaker estimation, thereby more effectively reducing the authentication time required for each user. Furthermore, because the authentication device P1 generates the second authentication word at the timing when the user is estimated in the first speaker estimation, even when another person (a third party) attempts to impersonate the user using recorded data of the user's speech or synthesized speech, it becomes difficult for the other person (a third party) to predict the second authentication word, and impersonation can be more effectively suppressed.
[0063] (Technology 3) the second authentication words include a plurality of words randomly selected from a word dictionary; A voice authentication method according to (Technology 1) or (Technology 2). With this configuration, the authentication device P1 can more effectively prevent others from impersonating the user by generating a second authentication word that includes a plurality of randomly selected words.
[0064] (Technology 4) the second authentication word is a sentence including a plurality of words randomly selected from a word dictionary; The voice authentication method according to any one of (Technology 1) to (Technology 3). With this configuration, the authentication device P1 can more effectively prevent others from impersonating the user by generating a second authentication word that includes a plurality of randomly selected words.
[0065] (Technology 5) The second authentication word is a sentence including a plurality of numbers related to the current date and time for authenticating the speaker. The voice authentication method according to any one of (Technology 1) to (Technology 3). With this configuration, the authentication device P1 can more effectively prevent others from impersonating the user by generating a second authentication word that includes a plurality of numbers related to the current date and time.
[0066] (Technology 6) The second authentication word is a sentence including a plurality of words related to the current location for authenticating the speaker. The voice authentication method according to any one of (Technology 1) to (Technology 3). With this configuration, the authentication device P1 can more effectively prevent others from impersonating the user by generating a second authentication word that includes a plurality of numbers related to the current location.
[0067] (Technology 7) the feature of the enrollment speaker is extracted from enrollment speech data acquired by having the enrollment speaker speak a registration word for each of the enrollment speakers, and is enrolled; The first authentication word is the enrollment word. The voice authentication method according to any one of (Technology 1) to (Technology 3). With this configuration, the authentication device P1 can perform more accurate voice authentication by setting the first authentication word to the same word as the registration word indicating pre-registered user information.
[0068] (Technology 8) an output unit (communication unit 10) that outputs a first authentication word to be spoken by a speaker (i.e., a user) and a second authentication word different from the first authentication word; an acquisition unit (communication unit 10) that acquires first utterance voice data obtained by having the speaker utter the first authentication word and second utterance voice data obtained by having the speaker utter the second authentication word; an estimation unit (processor 11) that performs a first speaker estimation to estimate which of the registered speakers the speaker is based on a first feature of the speaker extracted from the first speech voice data and a feature of at least one registered speaker registered in advance, and a second speaker estimation to estimate which of the registered speakers the speaker is based on a second feature of the speaker extracted from the second speech voice data and the feature of the registered speaker when it is determined that the speech content based on the second speech voice data matches the second authentication word; and an authentication unit (processor 11) that authenticates the speaker based on whether or not the enrolled speaker estimated by the first speaker estimation matches the enrolled speaker estimated by the second speaker estimation. Voice authentication device (authentication device P1). With this configuration, the authentication device P1 generates the second authentication word only when the user is estimated in the first speaker estimation, and performs the second speaker estimation, thereby more effectively reducing the authentication time required for each user. Furthermore, because the authentication device P1 generates the second authentication word at the timing when the user is estimated in the first speaker estimation, even when another person (a third party) attempts to impersonate the user using recorded data of the user's speech or synthesized speech, it becomes difficult for the other person (a third party) to predict the second authentication word, and impersonation can be more effectively suppressed.
[0069] Although various embodiments have been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the various embodiments described above may be combined in any manner without departing from the spirit of the invention. [Industrial Applicability]
[0070] The present disclosure is useful as a voice authentication method and a voice authentication device that more effectively prevent a decrease in voice authentication accuracy and spoofing by others in voice authentication. [Explanation of symbols]
[0071] 10. Communications Department 11 processors 12 Memory 100,100A Voice Authentication System 123 Word Dictionary DB Registered speaker database DP, DP1 Information display device DP2 display MC,MC1 Sound pickup device MC2 sound collection section P1 Authentication Device U1 User terminal
Claims
1. 1. A method of voice authentication performed by at least one computer, comprising: acquiring first speech data obtained by having a speaker speak a first authentication word; performing a first speaker estimation for estimating which of the registered speakers the speaker is, based on a first feature of the speaker extracted from the first speech voice data and a feature of at least one registered speaker who has been registered in advance; acquiring second speech data by having the speaker speak a second authentication word different from the first authentication word; if it is determined that the speech content based on the second speech voice data matches the second authentication word, a second speaker estimation is performed to estimate which of the registered speakers the speaker is, based on a second feature of the speaker extracted from the second speech voice data and the feature of the registered speaker; authenticating the speaker based on whether or not the enrollment speaker estimated by the first speaker estimation matches the enrollment speaker estimated by the second speaker estimation; Voice authentication method.
2. generating the second speech voice data when it is estimated by the first speaker estimation that the speaker is one of the enrolled speakers; The voice authentication method according to claim 1 .
3. the second authentication word includes a plurality of randomly selected words; The voice authentication method according to claim 1 .
4. the second authentication word is a sentence including a plurality of randomly selected words; The voice authentication method according to claim 1 .
5. the second authentication word is a sentence including a plurality of numbers related to a current date and time for authenticating the speaker; The voice authentication method according to claim 1 .
6. the second authentication word is a sentence including a plurality of words related to the current location for authenticating the speaker; The voice authentication method according to claim 1 .
7. the feature of the enrollment speaker is extracted from enrollment speech data acquired by having the enrollment speaker speak a registration word for each of the enrollment speakers, and is enrolled; the first authentication word is the enrollment word; The voice authentication method according to claim 1 .
8. an output unit that outputs a first authentication word to be spoken by a speaker and a second authentication word different from the first authentication word; an acquiring unit that acquires first uttered voice data obtained by having the speaker utter the first authentication word and second uttered voice data obtained by having the speaker utter the second authentication word; an estimation unit that performs a first speaker estimation to estimate which of the registered speakers the speaker is based on a first feature of the speaker extracted from the first speech voice data and a feature of at least one registered speaker registered in advance, and a second speaker estimation to estimate which of the registered speakers the speaker is based on the second feature of the speaker extracted from the second speech voice data and the feature of the registered speaker when it is determined that the speech content based on the second speech voice data matches the second authentication word; an authentication unit that authenticates the speaker based on whether or not an enrollment speaker estimated by the first speaker estimation matches an enrollment speaker estimated by the second speaker estimation. Voice authentication device.
Citation Information
Patent Citations
Information processing device and program
JP2021197060A