Personal authentication device

The personal authentication device generates a unique authentication sentence with specified speech characteristics and enforces a deadline, using voiceprint and speaking method verification to enhance authentication accuracy against voice synthesis and pre-recorded threats.

JP7726199B2Active Publication Date: 2025-08-20TOYOTA JIDOSHA KK
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022210666
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-08-20
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

Existing voice authentication systems are vulnerable to deception using pre-recorded voice data and voice synthesis technology, which cannot be effectively countered by existing methods.

Method used

A personal authentication device that generates a unique authentication sentence with specified speech characteristics, requiring the user to speak within a deadline, and authenticates based on voiceprint, wording, and speaking method matches to ensure accuracy.

Benefits of technology

Effectively counters deception using voice synthesis technology by ensuring the user's voice input matches the generated sentence, voiceprint, and speaking style, enhancing authentication accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007726199000001
    Figure 0007726199000001
  • Figure 0007726199000002
    Figure 0007726199000002
  • Figure 0007726199000003
    Figure 0007726199000003
Patent Text Reader

Abstract

To obtain an identity authentication device capable of dealing with deception using speech synthesis technology.SOLUTION: In an identity authentication device 100, an authentication text generated for authentication by a text generator 10 is presented to a user 50 by an authentication voice presentation unit 14 together with a speech method set by a speech method specification unit 12, and authentication based on the voiceprint, text, and speech method detected from voice data input to a voice input unit 16 is performed by a voiceprint authentication unit 18, voice authentication unit 20, speech method determination unit 22, and authentication result determination unit 26.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an authentication device that authenticates a user based on input voice data. [Background technology]

[0002] Voice authentication is used to authenticate a user based on their voice when entering a specific area or when granting permission to use an item, etc. However, there is a risk that authentication may be deceived by playing back recorded data of the registered user's voice or by using synthesized voices generated by speech synthesis technology, which has become more accurate in recent years.

[0003] Patent Document 1 discloses an invention in which a keyword extracted from received broadcast information such as that broadcast on the radio is presented to a user as an authentication phrase, the user is asked to speak the presented authentication phrase, and the user is authenticated based on the degree of match between the voice recognition of the speech and the authentication phrase. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2007-011965 Summary of the Invention [Problem to be solved by the invention]

[0005] Since pre-recorded voice data cannot handle randomly presented authentication phrases, the invention described in Patent Document 1 can eliminate deception using recorded voice data. However, there is a risk that the invention described in Patent Document 1 will not be able to handle synthetic voices that read out authentication phrases using voice synthesis technology.

[0006] In consideration of the above, an object of the present invention is to provide an individual authentication device that can deal with deception using voice synthesis technology. [Means for solving the problem]

[0007] To achieve the above objectives According to the first aspect The personal authentication device presents an authentication sentence generated for authentication together with a speech method set for authentication to a person to be authenticated, and performs authentication based on a voiceprint, wording, and speech method detected from input voice data. The person to be authenticated is authenticated when the degree of match between the wording detected from the input voice data and the authentication sentence is equal to or greater than a predetermined voice determination threshold, the degree of match between the voiceprint detected from the input voice data and the voiceprint of the registered user's voice data is equal to or greater than a predetermined voiceprint determination threshold, and the degree of match between each of the speaking speed, intonation, and emotional expression detected from the input voice data and each of the speaking speed, intonation, and emotional expression of the set speaking method is equal to or greater than a predetermined speaking method determination threshold.

[0008] According to the first aspect According to the personal authentication device, the authentication text generated during authentication is presented to the person to be authenticated along with the speech method set during authentication, thereby countering deception using pre-recorded voice data and voice synthesis technology. In addition, the person to be authenticated is authenticated only if the wording, voiceprint, and speaking method detected from the input voice data are correct, which improves the accuracy of authentication and makes it possible to counter deception using voice synthesis technology.

[0011] According to the second aspect The personal authentication device is In a first aspect, Authentication is performed using voice data entered within the authentication period.

[0012] According to the second aspect According to the personal authentication device, by setting an authentication deadline during which voice input is possible, it becomes difficult to generate voice using voice synthesis technology.

[0013] According to the third aspect The personal authentication device is In a second aspect, The authentication time limit is a time period from when the authentication text is presented to the authentication subject, and is approximately twice the time required for the subject to finish reading the authentication text at the presented speaking speed.

[0014] According to the third aspect According to the personal authentication device, by setting the authentication deadline to an appropriately short time, it is possible to counter deception using voice synthesis technology. [Effects of the Invention]

[0017] As described above, the personal authentication device according to the present invention can deal with deception using voice synthesis technology. [Brief explanation of the drawings]

[0018] [Figure 1]FIG. 2 is a block diagram showing an example of a specific configuration of the personal authentication device according to the present embodiment. [Figure 2] 10 is a flowchart showing an example of a process for generating authentication information such as an authentication text by the personal authentication device. [Figure 3] 10 is an example of an authentication voice prompt presented to a user. [Figure 4] 10 is a flowchart showing an example of an authentication voice generation process performed by the personal authentication device. [Figure 5] 10 is a flowchart showing an example of a personal authentication process performed by the personal authentication device. DETAILED DESCRIPTION OF THE INVENTION

[0019] Hereinafter, the personal authentication device 100 according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of a specific configuration of the personal authentication device 100 according to this embodiment. The personal authentication device 100 is a computer such as a PC that is constructed by applying machine learning, which will be described later. the authentication text; a voice input unit 16 that records the voice spoken by the user 50; a voiceprint authentication unit 18 that performs voiceprint analysis on the recorded voice data to authenticate whether the voice data belongs to a registered user; a voice authentication unit 20 that transcribes (converts to text) the recorded voice data using voice recognition and determines whether the transcription result matches the authentication text; a speaking method determination unit 22 that determines whether the speaking method of the recorded voice data matches the specified speaking method; an authentication deadline determination unit 24 that determines whether the voice data was input by the user 50 within a predetermined time; and an authentication result determination unit 26 that integrates the determination results of the voiceprint authentication unit 18, the voice authentication unit 20, the speaking method determination unit 22, and the authentication deadline determination unit 24 to make a final authentication determination.

[0020] In this embodiment, as an example, the voice input unit 16 may be a microphone built into the personal authentication device 100, or may be a mobile information terminal such as a smartphone. In such a case, an authentication application for performing voice authentication according to this embodiment is pre-installed on the smartphone.

[0021] 2 is a flowchart showing an example of a process for generating authentication information such as an authentication text by the personal authentication device 100. The process shown in FIG. 2 is executed when the user 50 starts authentication. In step S100, the text generation unit 10 generates an authentication text using a text generation AI. The text generation AI is constructed, for example, by having a machine learning model learn vocabulary by part of speech.

[0022] In step S102, the speech method specifying unit 12 generates a speech rate that specifies the speech rate for the generated authentication text. The speech rate is selected from three modes: slow, medium, and fast.

[0023] In step S104, the speech style specifying unit 12 generates an intonation that specifies the intonation (accent) to use when reading the generated authentication text. The intonation is specified to emphasize some of the words included in the authentication text. The number of words to be emphasized is about one to two, depending on the length of the authentication text.

[0024] In step S106, the speech style designation unit 12 performs emotion generation to designate an emotion (positive or negative) to be used when reading out the generated authentication text. The speech style designation unit 12 designates each of the speech style, ie, speech rate, intonation, and emotion, but may designate not only all of the speech rate, intonation, and emotion, but also some of the speech rate, intonation, and emotion.

[0025] In step S108, the authentication voice presentation unit 14 displays the authentication voice on the display of the mobile information terminal such as a smartphone of the user 50, and the process ends. FIG. 3 is an example of an authentication voice display presented to the user 50. As shown in FIG. 3, an authentication sentence 30 is displayed on the display. Along with the authentication sentence 30, the display also shows the speech method of "Emotion: Positive," check marks 34 indicating intonation, and arrows 32 indicating the speaking rate. The user 50 reads the presented authentication sentence 30 aloud by emphasizing the pronunciation of the words with check marks 34 and following the arrows 32 with the presented emotion.

[0026] In the process shown in FIG. 3, the authentication text 30 is presented to the user 50 and the user 50 is asked to speak, but it is also possible to play an authentication voice generated from the authentication text using voice synthesis technology on the mobile information terminal of the user 50 and have the user 50 speak in a way that imitates the authentication voice that has been played back.

[0027] Fig. 4 is a flowchart showing an example of authentication voice generation processing by the personal authentication device 100. In the processing shown in Fig. 4, the procedure of steps S100 to S106 is the same as steps S100 to S106 shown in Fig. 2, so the same reference numerals are used and detailed description will be omitted.

[0028] In step S208, speech synthesis technology is used to generate speech for authentication that reflects the words of the authentication sentence generated in step S100 at the speaking rate specified in step S102, with the intonation specified in step S104, and with the emotion specified in step S106.

[0029] In step S210, the generated authentication voice is played back on the mobile information terminal of the user 50, and the process ends.

[0030] 5 is a flowchart showing an example of identity authentication processing by identity authentication device 100. In step S300, user 50 speaks the presented authentication sentence according to the presented speaking method, and inputs the spoken voice into voice input unit 16. The voice data input into voice input unit 16 is input to voiceprint authentication unit 18, voice authentication unit 20, speaking method determination unit 22, and authentication deadline determination unit 24.

[0031] In step S302, it is determined whether the voice input in step S300 was performed within the authentication deadline. The authentication deadline is the time from when the authentication text is presented to the user 50, and in this embodiment, for example, is approximately twice the time it takes to finish reading the authentication text at the presented speaking speed. The requirement that the voice input be performed within the authentication deadline is, for example, a voice synthesis technology measure. This is because voice synthesis according to the specified speaking method cannot be performed accurately within a time approximately twice the time it takes to finish reading the authentication text at the presented speaking speed. If the voice input is performed within the authentication deadline in step S302, the procedure proceeds to step S304. If the voice input is not performed within the authentication deadline, the procedure proceeds to step S320.

[0032] In step S304, the voiceprint authentication unit 18 compares the input voice data with pre-registered voiceprint information of the user 50 and determines whether the voiceprint detected from the input voice data matches the pre-registered voiceprint information. If the voiceprint of the input voice data matches the pre-registered voiceprint information in step S304, the procedure proceeds to step S306. If the voiceprint of the input voice data does not match the pre-registered voiceprint information, the procedure proceeds to step S320. In this embodiment, the voiceprint authentication unit 18 is configured to analyze the voiceprint of the input voice data and calculate a match indicating the degree to which the voiceprint obtained by the analysis matches the voiceprint of the registered user's voice data using a trained model obtained by machine learning using voice data from an actual subject and voiceprint data as training data. The voiceprint authentication unit 18 determines whether the match is equal to or greater than a predetermined voiceprint determination threshold. The voiceprint determination threshold is set, for example, depending on the purpose of identity authentication. When strict security is required, the voiceprint determination threshold is set to a high value, and when strict security is not required, the voiceprint determination threshold is set to a low value.

[0033] In step S306, the voice authentication unit 20 transcribes the input voice data into a sentence (text) using voice recognition. Then, in step S308, the voice authentication unit 20 determines whether the sentence transcribed in step S306 matches the authentication sentence. If the transcribed sentence matches the authentication sentence in step S308, the procedure proceeds to step S310. If the transcribed sentence does not match the authentication sentence, the procedure proceeds to step S320. In this embodiment, the voice authentication unit 20 is configured to use a trained model obtained by machine learning using voice data from an actual subject and the sentence of the voice data as training data to convert the input voice data into a sentence according to the authentication sentence and to calculate a degree of match between the sentence and the authentication sentence. The voice authentication unit 20 determines whether the degree of match between the sentence and the authentication sentence is equal to or greater than a predetermined voice determination threshold. The voice determination threshold is set according to the purpose of personal authentication, as with voiceprint authentication by the voiceprint authentication unit 18.

[0034] In step S310, the speech style determination unit 22 calculates the speaking rate of the input speech data. In step S312, the speech style determination unit 22 calculates the intonation of the input speech data. Then, in step S314, the speech style determination unit 22 calculates the emotion of the input speech data. In this embodiment, the speech style determination unit 22 is configured to be able to calculate the speaking rate, intonation, and emotion of the input speech data using a trained model obtained by machine learning using speech data from an actual subject as training data, and to be able to calculate a degree of match indicating to what extent each feature of the speaking rate, intonation, and emotion matches a feature of a comparison target.

[0035] In step S316, the speech style determination unit 22 determines whether the features related to the speech style, such as the speech rate, intonation, and emotion indicated by the speech data, calculated in steps S310 to S314, match the features of the speech rate, intonation, and emotion specified by the speech style designation unit 12. In step S316, it determines whether the degree of agreement of each feature calculated by the speech style determination unit 22 is equal to or greater than a speech style determination threshold determined in advance for each feature. The speech style determination threshold is set depending on the purpose of personal authentication, as with voiceprint authentication by the voiceprint authentication unit 18.

[0036] In step S318, if the voice input is within the authentication deadline, the voiceprint of the input voice data matches the pre-registered voiceprint information, the text of the input voice matches the authentication sentence, and the feature amount related to the speaking method detected from the input voice data matches the feature amount related to the speaking method specified by the speaking method designation unit 12, the authentication result determination unit 26 makes a final authentication determination, authenticates that the user 50 is a registered user, and ends the process.

[0037] In step S320, since any one of the authentication period, voiceprint, wording, and speaking method does not satisfy the authentication requirements, the authentication result determination unit 26 determines that the user 50 is not a registered user and ends the process.

[0038] As described above, according to this embodiment, a newly generated authentication text is presented to the user 50, who is the subject of authentication, and the authentication text is input by having the user 50 speak it within the authentication deadline, and authentication is performed based on the voiceprint, wording, and speaking method of the input voice data, thereby eliminating deception using voice data generated by voice synthesis technology or recorded voice data.

[0039] When only an authentication sentence is presented and authentication is performed based on the input voice text and voiceprint, there is a risk that a malicious user could fraudulently bypass authentication by using authentication voice generated by voice synthesis technology. In this embodiment, by specifying the speech rate, intonation, and emotion when speaking the authentication sentence, it becomes difficult to synthesize voice within a short period of time, such as the authentication deadline, and fraudulent authentication can be prevented. [Explanation of symbols]

[0040] 10 Sentence generation section 12 Speech method specification section 14. Authentication voice prompt 16 Audio input section 18 Voiceprint Recognition Department 20 Voice authentication section 22 Speech method determination unit 24 Certification Expiration Determination Unit 26 Authentication result determination unit 30 Authentication Text 50 users 100 Personal authentication device

Claims

1. An authentication device that presents an authentication sentence generated for authentication together with a speech method set for authentication to a person to be authenticated, and performs authentication based on a voiceprint, a wording, and a speech method detected from input voice data, An identity authentication device that authenticates the person to be authenticated when the degree of match between the wording detected from the input voice data and the authentication sentence is equal to or greater than a predetermined voice determination threshold, the degree of match between the voiceprint detected from the input voice data and the voiceprint of the registered user's voice data is equal to or greater than a predetermined voiceprint determination threshold, and the degree of match between each of the speaking speed, intonation, and emotional expression detected from the input voice data and each of the speaking speed, intonation, and emotional expression of the set speaking method is equal to or greater than a predetermined speaking method determination threshold.

2. An identity authentication device as described in claim 1 that performs authentication using voice data entered within the authentication deadline.

3. An identity authentication device as described in Claim 2, wherein the authentication deadline is the time from when the authentication text is presented to the person to be authenticated, and is approximately twice the time it takes to finish reading the authentication text at the presented speaking speed.

Citation Information

Patent Citations

  • Method and device for speaker collating

    JP1996286692A

  • Individual collating device and recording medium with program for realizing it recorded thereon

    JP2001222295A

  • Speaker collating device and method

    JP2001265387A

  • System and method for personal authentication, and management system

    JP2007011965A

  • Device, method and program for discrimination of synthesized speech

    JP2010237364A