Authentication device, authentication method, and recording medium
The authentication device and method enhance voice recognition accuracy by combining air and bone conduction feature amounts and calculating a difference feature based on frequency spectra, effectively addressing limitations in traditional authentication technologies.
Patent Information
- Application Number
- JP2023546610
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-09-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-09-08
AI Technical Summary
Existing authentication technologies using voice recognition face challenges in accurately authenticating subjects, particularly in noisy environments or when traditional authentication methods like face or fingerprint recognition are compromised, such as with masks or gloves.
The proposed solution involves an authentication device and method that calculates air conduction and bone conduction feature amounts from voice signals, combining these to create a target feature amount for more accurate authentication. Additionally, a difference feature amount is calculated based on the frequency spectra of air and bone conduction signals, enhancing authentication accuracy by accounting for skeletal characteristics.
This approach improves authentication accuracy by utilizing both air and bone conduction voice signals, reducing processing load, and effectively addressing limitations in traditional authentication methods, especially in challenging environments.
Smart Images

Figure 0007697517000001 
Figure 0007697517000002 
Figure 0007697517000003
Abstract
Description
Technical Field
[0001] This disclosure relates to the technical field of authentication devices, authentication methods, and recording media that can authenticate a subject using, for example, the voice of the subject.
Background Art
[0002] An example of an authentication device that can authenticate a subject using the voice of the subject is described in Patent Document 1.
[0003] In addition, Patent Documents 2 to 4 are cited as prior art documents related to this disclosure.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0005] This disclosure aims to provide an authentication device, an authentication method, and a recording medium for improving the technologies described in the prior art documents.
Means for Solving the Problems
[0006] The first aspect of the authentication device calculates an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a bone conduction feature amount, which is a feature amount of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of the subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject, and calculates a target feature amount, which is a feature amount of the voice of the subject, by combining the air conduction feature amount and the bone conduction feature amount, and includes a calculation means for calculating the target feature amount and an authentication means for authenticating the subject based on the target feature amount.
[0007] The second aspect of the authentication device calculates an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a difference feature amount, which is a feature amount of the difference between the frequency spectrum of the air conduction voice signal and the frequency spectrum of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of the subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject, and includes a calculation means for calculating the air conduction feature amount and the difference feature amount and an authentication means for authenticating the subject based on the air conduction feature amount and the difference feature amount.
[0008] The first aspect of the authentication method calculates an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a bone conduction feature amount, which is a feature amount of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of the subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject, calculates a target feature amount, which is a feature amount of the voice of the subject, by combining the air conduction feature amount and the bone conduction feature amount, and authenticates the subject based on the target feature amount.
[0009] The second aspect of the authentication method calculates an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a difference feature amount, which is a feature amount of the difference between the frequency spectrum of the air conduction voice signal and the frequency spectrum of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of the subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject, and authenticates the subject based on the air conduction feature amount and the difference feature amount.
[0010] The first aspect of the recording medium is a computer program recorded on a recording medium that causes a computer to calculate an air conduction feature amount, which is a feature amount of an air conduction voice signal indicating an air conduction sound of a target person's voice, and a bone conduction feature amount, which is a feature amount of a bone conduction voice signal indicating a bone conduction sound of the target person's voice, from the air conduction voice signal and the bone conduction voice signal of the target person's voice, and to calculate a target feature amount, which is a feature amount of the target person's voice, by combining the air conduction feature amount and the bone conduction feature amount, and to execute an authentication method for authenticating the target person based on the target feature amount.
[0011] The second aspect of the recording medium is a computer program recorded on a recording medium that causes a computer to calculate an air conduction feature amount, which is a feature amount of an air conduction voice signal indicating an air conduction sound of a target person's voice, and a difference feature amount, which is a feature amount of the difference between the frequency spectrum of the air conduction voice signal and the frequency spectrum of the bone conduction voice signal, from the air conduction voice signal and the bone conduction voice signal of the target person's voice, and to execute an authentication method for authenticating the target person based on the air conduction feature amount and the difference feature amount.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0013] Hereinafter, embodiments of an authentication device, an authentication method, and a recording medium will be described with reference to the drawings.
[0014] (1) First Embodiment First, a first embodiment of an authentication device, an authentication method, and a recording medium will be described. Hereinafter, the first embodiment of the authentication device, the authentication method, and the recording medium will be described using an authentication device 1000 to which the first embodiment of the authentication device, the authentication method, and the recording medium is applied.
[0015] FIG. 1 is a block diagram showing the configuration of the authentication device 1000 in the first embodiment. As shown in FIG. 1, the authentication device 1000 includes a calculation unit 1001 and an authentication unit 1002.
[0016] In the first example, the calculation unit 1001 calculates an air conduction feature amount, which is a feature amount of an air conduction audio signal, from an air conduction audio signal indicating the air conduction sound of the voice of the subject (that is, the voice spoken by the subject, the same hereinafter). Further, the calculation unit 1001 calculates a bone conduction feature amount, which is a feature amount of the bone conduction audio signal, from a bone conduction audio signal indicating the bone conduction sound of the voice of the subject. Further, the calculation unit 1001 calculates a target feature amount, which is a feature amount of the subject, by combining the air conduction audio signal and the bone conduction feature amount. The authentication unit 1002 authenticates the subject based on the target feature amount calculated by the calculation unit 1001.
[0017] As described above, in the first example, the authentication device 1000 authenticates the subject based on not only the air conduction feature amount indicating the features of the subject's voice itself but also the bone conduction feature amount indicating the features of the subject's voice with the influence of the subject's skeleton superimposed thereon (that is, the bone conduction feature amount also indicating the features of the subject's skeleton). Therefore, compared with an authentication device that authenticates the subject based on either the air conduction feature amount or the bone conduction feature amount, the authentication device 1000 can authenticate the subject more accurately using the subject's voice. In particular, the authentication device 1000 does not have to separately perform the process of authenticating the subject based on the air conduction feature amount and the process of authenticating the subject based on the bone conduction feature amount different from the air conduction feature amount. That is, the authentication device 1000 only has to perform the process of authenticating the subject based on the target feature amount calculated from the combined air conduction feature amount and bone conduction feature amount. Therefore, the authentication device 1000 can reduce the processing load for authenticating the subject.
[0018] On the other hand, in the second example, the calculation unit 1001 calculates a difference feature amount, which is a feature amount of the difference between the frequency spectrum of the air conduction voice signal indicating the air conduction sound of the subject's voice and the frequency spectrum of the bone conduction voice signal indicating the bone conduction sound of the subject's voice, from the air conduction voice signal and the bone conduction voice signal of the subject's voice. Further, the calculation unit 1001 calculates an air conduction feature amount, which is a feature amount of the air conduction voice signal, from the air conduction voice signal. The authentication unit 1002 authenticates the subject based on the air conduction feature amount and the difference feature amount.
[0019] Here, as described above, the air conduction feature amount indicates the feature of the subject's voice itself. Further, the difference feature amount corresponds to a feature amount in which the feature of the subject's voice itself is substantially excluded from the feature of the subject's voice with the influence of the subject's skeleton superimposed thereon. That is, the difference feature amount corresponds to a feature amount indicating the feature of the subject's skeleton (that is, the skeleton unique to the subject) indicating the individuality of the subject. For this reason, the authentication device 1000 authenticates the subject based on the air conduction feature amount indicating the feature of the subject's voice itself and the difference feature amount indicating the feature of the subject's skeleton itself. As a result, compared with an authentication device that authenticates a subject based on either one of the air conduction feature amount and the difference feature amount, the authentication device 1000 can authenticate the subject with higher accuracy using the subject's voice.
[0020] (2) Second Embodiment Subsequently, a second embodiment of the authentication device, the authentication method, and the recording medium will be described. Hereinafter, the second embodiment of the authentication device, the authentication method, and the recording medium will be described using an authentication system SYS to which the second embodiment of the authentication device, the authentication method, and the recording medium is applied.
[0021] (2-1) Configuration of Authentication System SYS First, with reference to FIG. 2, the configuration of the authentication system SYS in the second embodiment will be described. FIG. 2 is a block diagram showing the configuration of the authentication system SYS in the second embodiment.
[0022] As shown in FIG. 2, the authentication system SYS includes an air conduction microphone 1, a bone conduction microphone 2, and an authentication device 3.
[0023] The air-conduction microphone 1 is a voice detection device capable of detecting the air-conducted sound of the voice of a subject. Specifically, it detects the air-conducted sound of the subject's voice by detecting the vibration of the air generated along with the subject's voice. The air-conduction microphone 1 generates a voice signal indicating the air-conducted sound by detecting the air-conducted sound. In the following description, the voice signal indicating the air-conducted sound is referred to as the "air-conducted voice signal". The air-conduction microphone 1 outputs the generated air-conducted voice signal to the authentication device 3.
[0024] The bone-conduction microphone 2 is a voice detection device capable of detecting the bone-conducted sound of the voice of a subject. Specifically, it detects the bone-conducted sound of the subject's voice by detecting the vibration of the subject's bone (skeleton) generated along with the subject's voice. The bone-conduction microphone 2 generates a voice signal indicating the bone-conducted sound by detecting the bone-conducted sound. In the following description, the voice signal indicating the bone-conducted sound is referred to as the "bone-conducted voice signal". The bone-conduction microphone 2 outputs the generated bone-conducted voice signal to the authentication device 3.
[0025] The authentication device 3 performs an authentication operation for authenticating the subject using the voice of the subject. That is, the authentication device 3 performs voice authentication. To perform the authentication operation, the authentication device 3 acquires the air-conducted voice signal from the air-conduction microphone 1. Further, the authentication device 3 acquires the bone-conducted voice signal from the bone-conduction microphone 2. Thereafter, the authentication device 3 authenticates the subject using the air-conducted voice signal and the bone-conducted voice signal.
[0026] A device including the air-conduction microphone 1, the bone-conduction microphone 2, and the authentication device 3 may be used as the authentication system SYS. For example, a mobile terminal (e.g., a smartphone) including the air-conduction microphone 1 and the bone-conduction microphone 2 and capable of functioning as the authentication device 3 may be used as the authentication system SYS. For example, a wearable device including the air-conduction microphone 1, the bone-conduction microphone 2, and the authentication device 3 may be used as the authentication system SYS.
[0027] As an example of a scenario where the authentication system SYS for voice authentication is applied, there is a scenario where it is not easy to accurately perform face authentication and iris authentication. As an example of a scenario where it is not easy to accurately perform face authentication and iris authentication, there is a scenario of authenticating a subject wearing a mask. For example, the authentication system SYS may be used to manage the entry of workers wearing masks at at least one of a construction site and a factory. For example, the authentication system SYS may be used to manage the entry and exit of medical staff wearing masks in a medical facility. As another example of a scenario where the authentication system SYS for voice authentication is applied, there is a scenario where it is not easy to accurately perform fingerprint authentication. As an example of a scenario where it is not easy to accurately perform fingerprint authentication, there is a scenario of authenticating a subject wearing gloves. For example, the authentication system SYS may be used to manage the entry and exit of medical staff wearing gloves in a medical facility. As another example of a scenario where the authentication system SYS for voice authentication is applied, there is a scenario of authenticating a subject via a telephone service. However, the scenarios where the authentication system SYS is applied are not limited to the scenarios described here.
[0028] (2-2) Configuration of Authentication Device 3 Subsequently, with reference to FIG. 3, the configuration of the authentication device 3 in the second embodiment will be described. FIG. 3 is a block diagram showing the configuration of the authentication device 3 in the second embodiment.
[0029] As shown in FIG. 3, the authentication device 3 includes an arithmetic device 31 and a storage device 32. Further, the authentication device 3 may include a communication device 33, an input device 34, and an output device 35. However, the authentication device 3 may not include at least one of the communication device 33, the input device 34, and the output device 35. The arithmetic device 31, the storage device 32, the communication device 33, the input device 34, and the output device 35 may be connected via a data bus 36.
[0030] The arithmetic unit 31 includes, for example, at least one of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and an FPGA (Field Programmable Gate Array). The arithmetic unit 31 reads a computer program. For example, the arithmetic unit 31 may read the computer program stored in the storage device 32. For example, the arithmetic unit 31 may read the computer program stored in a computer-readable and non-transitory recording medium using a recording medium reading device (for example, the input device 34 described later) provided in the authentication device 3. The arithmetic unit 31 may obtain a computer program from a device (not shown) disposed outside the authentication device 3 via the communication device 33 (or other communication devices) (that is, it may be downloaded or read). The arithmetic unit 31 executes the read computer program. As a result, logical functional blocks for executing the operations (for example, the authentication operation described above) to be performed by the authentication device 3 are realized within the arithmetic unit 31. That is, the arithmetic unit 31 can function as a controller for realizing logical functional blocks for executing the operations (in other words, processes) to be performed by the authentication device 3.
[0031] FIG. 3 shows an example of the logical functional blocks realized in the arithmetic unit 31 for executing the authentication operation. As shown in FIG. 3, a calculation unit 311, which is a specific example of a “calculation means”, and an authentication unit 312, which is a specific example of an “authentication means”, are realized within the arithmetic unit 31.
[0032] The calculation unit 311 calculates a target feature amount, which is a feature amount of the target person used in the authentication operation, from the air-conducted voice signal and the bone-conducted voice signal. Note that the target feature amount calculated by the calculation unit 311 will be described in detail later.
[0033] The authentication unit 312 authenticates the target person based on the target feature amount calculated by the calculation unit 311. That is, the authentication unit 312 determines whether the target person matches the registered person based on the target feature amount calculated by the calculation unit 311. Specifically, the registered feature amount, which is a feature amount related to the voice of the registered person, is pre-registered in the collation DB (DataBase) 321 stored in the storage device 32. In the collation DB 321, such registered feature amounts are registered for the number of registered persons. The authentication unit 312 determines whether the target person matches the registered person by comparing the target feature amount calculated by the calculation unit 311 with the registered feature amount registered in the collation DB 321.
[0034] The storage device 32 can store desired data. For example, the storage device 32 may temporarily store the computer program executed by the arithmetic unit 31. The storage device 32 may temporarily store the data temporarily used by the arithmetic unit 31 when the arithmetic unit 31 is executing the computer program. The storage device 32 may store the data that the authentication device 3 stores in the long term. Note that the storage device 32 may include at least one of a RAM (Random Access Memory), a ROM (Read Only Memory), a hard disk device, a magneto-optical disk device, an SSD (Solid State Drive), and a disk array device. That is, the storage device 32 may include a non-temporary recording medium.
[0035] The communication device 33 can communicate with a device outside the authentication device 3 via a communication network (not shown). For example, the communication device 33 may be able to communicate with at least one of the air conduction microphone 1 and the bone conduction microphone 2. In this case, the communication device 33 may receive (that is, acquire) an air conduction voice signal from the air conduction microphone 1 via a communication network (not shown). The communication device 33 may receive (that is, acquire) a bone conduction voice signal from the bone conduction microphone 2 via a communication network (not shown).
[0036] The input device 34 is a device that receives the input of information to the authentication device 3 from the outside of the authentication device 3. For example, the input device 34 may include an operating device (for example, at least one of a keyboard, a mouse, and a touch panel) that can be operated by an operator of the authentication device 3. For example, the input device 34 may include a reading device that can read information recorded as data on an externally attachable recording medium with respect to the authentication device 3. For example, the input device 34 may include an input interface to which at least one of an air conduction voice signal output from the air conduction microphone 1 and a bone conduction voice signal output from the bone conduction microphone 2 is input.
[0037] The output device 35 is a device that outputs information to the outside of the authentication device 3. For example, the output device 35 may output the information as an image. That is, the output device 35 may include a display device (so-called display) that can display an image indicating the information to be output. For example, the output device 35 may output the information as voice. That is, the output device 35 may include a voice device (so-called speaker) that can output voice. For example, the output device 35 may output the information on paper. That is, the output device 35 may include a printing device (so-called printer) that can print desired information on paper.
[0038] (2-3) Operation of Authentication Device 3 (Authentication Operation) Subsequently, the flow of the authentication operation performed by the authentication device 3 in the second embodiment will be described. In the second embodiment, the authentication device 3 performs at least one of a first authentication operation and a second authentication operation. Therefore, hereinafter, the first authentication operation and the second authentication operation will be described in order.
[0039] (2-3-1) First Authentication Operation First, with reference to FIG. 4, the flow of the first authentication operation performed by the authentication device 3 in the second embodiment will be described. FIG. 4 is a flowchart showing the flow of the first authentication operation performed by the authentication device 3 in the second embodiment.
[0040] As shown in FIG. 4, the calculation unit 311 acquires an air-conducted voice signal indicating the air-conducted sound of the voice of the subject from the air-conducted microphone 1 (step S11). Further, the calculation unit 311 acquires a bone-conducted voice signal indicating the bone-conducted sound of the voice of the subject from the bone-conducted microphone 2 (step S12).
[0041] Thereafter, the calculation unit 311 calculates an air-conducted feature amount, which is a feature amount of the air-conducted voice signal, from the air-conducted voice signal acquired in step S11 (step S13). Further, the calculation unit 311 calculates a bone-conducted feature amount, which is a feature amount of the bone-conducted voice signal, from the bone-conducted voice signal acquired in step S12 (step S13).
[0042] The calculation unit 311 may calculate, as the air-conducted feature amount, any parameter that qualitatively and / or quantitatively indicates the feature of the air-conducted voice signal. For example, the calculation unit 311 may calculate, as the air-conducted feature amount, any parameter that indicates the feature of the air-conducted voice signal by performing a desired voice analysis process on the air-conducted voice signal. As an example of the desired voice analysis process, at least one of a frequency analysis process, a cepstrum analysis process, and a pitch extraction process can be mentioned. As an example of any parameter that indicates the feature of the air-conducted voice signal, there is a mel frequency cepstrum coefficient (MFCC) that can be calculated from the result of the frequency analysis process performed on the air-conducted voice signal.
[0043] The air-conducted feature amount is an N-dimensional vector (that is, a vector composed of N vector elements). Note that "N" is a constant indicating an integer of 1 or more. In this case, the number of dimensions of the vector is preferably set to an appropriate number that enables the authentication operation to be performed appropriately. As an example, when the mel frequency cepstrum coefficient is used as the air-conducted feature amount, the air-conducted feature amount may be a vector of 12 dimensions or more.
[0044] Similarly, the calculation unit 311 may calculate, as bone conduction feature amounts, any parameters that qualitatively and / or quantitatively indicate the features of the bone conduction voice signal. For example, the calculation unit 311 may calculate, as bone conduction feature amounts, any parameters that indicate the features of the bone conduction voice signal by performing desired voice analysis processing on the bone conduction voice signal. As an example of any parameter that indicates the features of the bone conduction voice signal, there is the mel-frequency cepstral coefficient that can be calculated from the result of the frequency analysis processing performed on the bone conduction voice signal.
[0045] The bone conduction feature amount is an M-dimensional vector (that is, a vector composed of M vector elements). Note that "M" is a constant indicating an integer of 1 or more. In this case, it is preferable that the number of dimensions of the vector is set to an appropriate number that enables the authentication operation to be performed appropriately. As an example, when the mel-frequency cepstral coefficient is used as the air conduction feature amount, the bone conduction feature amount may be a vector of 12 dimensions or more.
[0046] Thereafter, the calculation unit 311 combines (in other words, concatenates or synthesizes) the air conduction feature amount calculated in step S13 and the bone conduction feature amount calculated in step S13 (step S14). As a result, the calculation unit 311 calculates a combined feature amount that is a feature amount composed of the combined air conduction feature amount and bone conduction feature amount (step S14).
[0047] As described above, since the air conduction feature amount is an N-dimensional vector and the bone conduction feature amount is an M-dimensional vector, the combined feature amount typically becomes an (N + M)-dimensional vector. That is, the number of dimensions of the combined feature amount is N + M. Conversely, the calculation unit 311 may calculate the combined feature amount so that the combined feature amount includes the N vector elements included in the air conduction feature amount and the M vector elements included in the bone conduction feature amount.
[0048] However, the combined feature quantity may be a vector with less than N+M dimensions. That is, the number of dimensions of the combined feature quantity may be less than N+M. However, the number of dimensions of the combined feature quantity is greater than N and greater than M. That is, the combined feature quantity may be a vector that is less than N+M dimensions, greater than N dimensions, and greater than M dimensions. As an example, the calculation unit 311 calculates the combined feature quantity such that the combined feature quantity includes at least one of N' vector elements (where N' is a constant indicating an integer of 1 or more and less than N) among the N vector elements included in the air conduction feature quantity and at least one of M' vector elements (where M' is a constant indicating an integer of 1 or more and less than M) among the M vector elements included in the bone conduction feature quantity. That is, the operation of "calculating the combined feature quantity by combining the air conduction feature quantity and the bone conduction feature quantity" in the second embodiment may mean the operation of "calculating the combined feature quantity such that the combined feature quantity includes at least one of the N vector elements included in the air conduction feature quantity and at least one of the M vector elements included in the bone conduction feature quantity".
[0049] Thereafter, the calculation unit 311 calculates a target feature quantity used by the authentication unit 312 for performing the authentication operation from the combined feature quantity calculated in step S14 (step S15). For example, the calculation unit 311 may calculate a target feature quantity corresponding to the extracted feature quantity by extracting a feature quantity indicating the features of the target person from the combined feature quantity calculated in step S14.
[0050] The calculation unit 311 may calculate the target feature quantity from the combined feature quantity by using a neural network that can output the target feature quantity when the combined feature quantity is input and can be constructed by machine learning. The neural network may be constructed in advance by machine learning using teacher data including an air conduction voice signal of a sample person, a bone conduction voice signal of the sample person, and a correct label of the authentication result of the sample person.
[0051] Thereafter, the authentication unit 312 authenticates the target person based on the target feature amount calculated in step S15 (step S16). Specifically, the authentication unit 312 calculates the similarity between the target feature amount calculated in step S15 and the registered feature amount corresponding to the registered person registered in the collation DB 321. When the calculated similarity exceeds a predetermined authentication threshold (that is, the target feature amount is similar to the registered feature amount), the authentication unit 312 may determine that the target person matches the registered person. On the other hand, when the calculated similarity is below the predetermined authentication threshold (that is, the target feature amount is not similar to the registered feature amount), the authentication unit 312 may determine that the target person does not match the registered person.
[0052] The authentication unit 312 may calculate the similarity using any method for calculating the similarity between two feature amounts. As an arbitrary method for calculating the similarity between two feature amounts, a method using a Probabilistic Linera Discriminant Analysis (PLDA) model can be mentioned.
[0053] The authentication unit 312 may authenticate the target person using a neural network. For example, the authentication unit 312 may authenticate the target person using a neural network to which a probabilistic linear discriminant analysis model is applied. The neural network may be pre-constructed by machine learning using teacher data including an air-conducted voice signal of a sample person, a bone-conducted voice signal of the sample person, and a correct label of the authentication result of the sample person.
[0054] As described above, when the calculation unit 311 uses a neural network, the neural network used by the calculation unit 311 and the neural network used by the authentication unit 312 may be integrated. That is, the calculation unit 311 may calculate the target feature amount using the first network part of the neural network, and the authentication unit 312 may authenticate the target person using the second network part of the neural network to which the output of the first network part is input. In this case, the neural networks used by the calculation unit 311 and the authentication unit 312 may be neural networks compliant with a method called so-called x-vector (in other words, Deep Speaker Embedding).
[0055] A plurality of registered feature amounts corresponding to a plurality of registered persons may be registered in the collation DB 321. In this case, the authentication unit 312 may repeatedly perform an operation of determining whether or not the target person matches a certain registered person by calculating the similarity between a certain registered feature amount corresponding to a certain registered person and the target feature amount from the collation DB 321 using the plurality of registered feature amounts.
[0056] When the first authentication operation is performed, the registered feature amounts registered in the collation DB 321 may be generated in the same flow as the target feature amounts used in the first authentication operation. Specifically, in order to register the registered feature amounts in the collation DB 321, first, an air conduction voice signal indicating the air conduction sound of the voice of the registered person and a bone conduction voice signal indicating the bone conduction sound of the voice of the registered person may be acquired. Thereafter, an air conduction feature amount may be calculated from the air conduction voice signal, and a bone conduction feature amount may be calculated from the bone conduction voice signal. Thereafter, a combined feature amount may be calculated by combining the air conduction feature amount and the bone conduction feature amount. Thereafter, a registered feature amount may be calculated from the combined feature amount.
[0057] When the first authentication operation is performed in the flow shown in FIG. 4 like this, the calculation unit 311 may include the functional blocks shown in FIG. 5. Specifically, as shown in FIG. 5, the calculation unit 311 may include a calculation unit 3111, a calculation unit 3112, a calculation unit 3113, and a calculation unit 3114. The calculation unit 3111 may calculate an air-conducted feature amount from the air-conducted voice signal. The calculation unit 3112 may calculate a bone-conducted feature amount from the bone-conducted voice signal. The calculation unit 3112 may calculate a combined feature amount by combining the air-conducted feature amount calculated by the calculation unit 3111 and the bone-conducted feature amount calculated by the calculation unit 3112. The calculation unit 3114 may calculate a target feature amount from the combined feature amount calculated by the calculation unit 3113.
[0058] According to the first authentication operation described above, the authentication device 3 authenticates the subject based on not only the air conduction feature amount indicating the feature of the subject's voice itself but also the bone conduction feature amount indicating the feature of the subject's voice with the influence of the subject's skeleton superimposed thereon (that is, the bone conduction feature amount also indicating the feature of the subject's skeleton). That is, the authentication device 3 authenticates the subject using both the air conduction voice signal and the bone conduction voice signal. As a result, compared with the authentication device of the first comparative example that authenticates the subject based on either the air conduction feature amount or the bone conduction feature amount (that is, authenticates the subject based on either the air conduction voice signal or the bone conduction voice signal), the authentication device 3 can authenticate the subject with higher accuracy using the subject's voice. This is because when the authentication device of the first comparative example authenticates the subject based on the air conduction feature amount (that is, does not use the bone conduction feature amount to authenticate the subject), there may arise a technical problem that the authentication accuracy may deteriorate when the acquisition environment of the air conduction voice signal is not appropriate. For example, the authentication accuracy may deteriorate when the acquisition environment of the air conduction voice signal is a noisy environment or an environment where the subject is not uttering the voice appropriately. On the other hand, when the authentication device of the first comparative example authenticates the subject based on the bone conduction feature amount (that is, does not use the air conduction feature amount to authenticate the subject), there may arise a technical problem that the authentication accuracy may deteriorate because the accuracy of the bone conduction voice signal is originally lower than that of the air conduction voice signal. However, in the first authentication operation, the authentication device 3 authenticates the subject based on both the air conduction feature amount and the bone conduction feature amount. Therefore, the authentication device 3 can appropriately solve the technical problems that may occur in the authentication device of the first comparative example.
[0059] Furthermore, according to the first authentication operation, the authentication device 3 does not necessarily perform separately the process of authenticating the subject based on the air conduction feature amount and the process of authenticating the subject based on the bone conduction feature amount different from the air conduction feature amount. That is, the authentication device 3 does not necessarily perform separately the two processes of authenticating the subject based on two different feature amounts. In other words, the authentication device 3 only needs to perform the process of authenticating the subject based on one type of feature amount called the target feature amount. Therefore, compared with the authentication device of the second comparative example that needs to separately perform the process of authenticating the subject based on the air conduction feature amount and the process of authenticating the subject based on the bone conduction feature amount, the authentication device 3 can reduce the number of times of performing the process of authenticating the subject based on the feature amount (for example, the number of times of calculating the similarity described above). As an example, the authentication device 3 can reduce the number of times of performing the process of authenticating the subject based on the feature amount to about half of the number of times of performing the process of authenticating the subject based on the feature amount by the authentication device of the second comparative example. As a result, the authentication device 3 can reduce the processing load for authenticating the subject.
[0060] Also, the authentication device 3 can calculate the target feature amount from the combined feature amount using a neural network. Therefore, even when a combined feature amount with a larger number of elements is used compared to each of the air conduction feature amount and the bone conduction feature amount, the authentication device 3 can relatively easily calculate the target feature amount.
[0061] (2-3-2) Second Authentication Operation Subsequently, with reference to FIG. 6, the flow of the second authentication operation performed by the authentication device 3 in the second embodiment will be described. FIG. 6 is a flowchart showing the flow of the second authentication operation performed by the authentication device 3 in the second embodiment.
[0062] As shown in FIG. 6, also in the second authentication operation, similar to the first authentication operation, the calculation unit 311 acquires an air conduction voice signal from the air conduction microphone 1 (step S11). Further, the calculation unit 311 acquires a bone conduction voice signal from the bone conduction microphone 2 (step S12).
[0063] After that, also in the second authentication operation, similarly to the first authentication operation, the calculation unit 311 calculates an air-conducted feature amount from the air-conducted voice signal acquired in step S11 (step S23).
[0064] On the other hand, in the second authentication operation, the calculation unit 311 does not necessarily calculate a bone-conducted feature amount from the bone-conducted voice signal acquired in step S12. In the second authentication operation, the calculation unit 311 calculates a difference feature amount instead of the bone-conducted feature amount (step S24). The difference feature amount is a feature amount indicating the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal (that is, a feature amount indicating the feature of the difference). For example, the difference itself between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal may be used as the difference feature amount. For example, a parameter calculated from the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal may be used as the difference feature amount. For example, a parameter that quantitatively or qualitatively indicates the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal may be used as the difference feature amount.
[0065] After that, the authentication unit 312 authenticates the subject based on the air-conducted feature amount calculated in step S23 (step S25). Further, the authentication unit 312 authenticates the subject based on the difference feature amount calculated in step S24 (step S26). Therefore, in the second embodiment, each of the air-conducted feature amount and the difference feature amount is used as a target feature amount actually used for authenticating the subject.
[0066] Also in the second authentication operation, similar to the first authentication operation, the authentication unit 312 authenticates the target person by calculating the similarity between the target feature amount and the registered feature amount registered in the collation DB 321. Here, as described above, in the second embodiment, each of the air conduction feature amount and the difference feature amount is used as the target feature amount. For this reason, in the second authentication operation, in the collation DB 321, as the registered feature amount, a first registered feature amount corresponding to the air conduction feature amount and a second registered feature amount corresponding to the difference feature amount are registered. The first registered feature amount is a feature amount of an air conduction voice signal indicating the air conduction sound of the voice of the registered person. The second registered feature amount is a feature amount indicating the difference between the frequency spectrum of the air conduction voice signal indicating the air conduction sound of the voice of the registered person and the frequency spectrum of the bone conduction voice signal indicating the bone conduction sound of the voice of the registered person. In this case, in step S25, the authentication unit 312 authenticates the target person by calculating the similarity between the air conduction feature amount calculated as the difference feature amount in step S23 and the first registered feature amount registered in the collation DB 321. Further, in step S26, the authentication unit 312 authenticates the target person by calculating the similarity between the difference feature amount calculated as the difference feature amount in step S24 and the second registered feature amount registered in the collation DB 321.
[0067] Thereafter, the authentication unit 312 authenticates the target person based on the authentication result of the target person in step S25 and the authentication result of the target person in step S26 (step S27). That is, in the second authentication operation, the authentication unit 312 tentatively authenticates the target person in each of steps S25 and S26, and in step S27, authenticates the target person definitely (in other words, finally) based on the tentative authentication result of the target person. As an example, when it is determined that the target person matches one registered person in step S25 and the target person matches the same one registered person in step S26, the authentication unit 312 may determine that the target person matches one registered person. On the other hand, when it is determined that the target person does not match one registered person in at least one of steps S25 and S26, the authentication unit 312 may determine that the target person does not match one registered person.
[0068] When the second authentication operation is performed in the flow shown in FIG. 6, the calculation unit 311 and the authentication unit 312 may have the functional blocks shown in FIG. 7. Specifically, as shown in FIG. 7, the calculation unit 311 may include the calculation unit 3111 shown in FIG. 5 and the calculation unit 3115. The authentication unit 312 may include the authentication unit 3121, the authentication unit 3122, and the authentication unit 3123. As described above, the calculation unit 3111 may calculate an air-conducted feature amount from the air-conducted voice signal. The calculation unit 3115 may calculate a differential feature amount from the air-conducted voice signal and the bone-conducted voice signal. The authentication unit 3121 may tentatively authenticate the subject based on the air-conducted feature amount calculated by the calculation unit 3111. The authentication unit 3122 may tentatively authenticate the subject based on the differential feature amount calculated by the calculation unit 3115. The authentication unit 3123 may definitely authenticate the subject based on the authentication result by the authentication unit 3121 and the authentication result by the authentication unit 3122.
[0069] According to the second authentication operation described above, similar to the first authentication operation, the authentication device 3 authenticates the subject using both the air-conducted voice signal and the bone-conducted voice signal. As a result, compared with the authentication device of the first comparative example that authenticates the subject based on either the air-conducted voice signal or the bone-conducted voice signal, the authentication device 3 can authenticate the subject with higher accuracy using the voice of the subject.
[0070] Furthermore, according to the second authentication operation, the authentication device 3 authenticates the subject based on the differential feature amount instead of the bone-conducted feature amount. Here, the differential feature amount corresponds to a feature amount in which the features of the subject's voice itself are substantially excluded from the features of the subject's voice with the influence of the subject's skeleton superimposed. That is, the differential feature amount corresponds to a feature amount indicating the features of the subject's skeleton (that is, the skeleton unique to the subject) that indicates the individuality of the subject. Therefore, the authentication device 3 authenticates the subject based on the air-conducted feature amount indicating the features of the subject's voice itself and the differential feature amount indicating the features of the subject's skeleton itself. As a result, compared with the authentication device of the third comparative example that authenticates the subject based on either the air-conducted feature amount or the differential feature amount, the authentication device 3 can authenticate the subject with higher accuracy using the voice of the subject.
[0071] Furthermore, the authentication device 3 authenticates the subject definitively based on the provisional authentication result of the subject based on each of the air conduction feature amount and the differential feature amount. For this reason, compared with the case where the authentication result of the subject based on the air conduction feature amount is used as it is as the definitive authentication result of the subject or the authentication result of the subject based on the differential feature amount is used as it is as the definitive authentication result of the subject, the authentication device 3 can authenticate the subject with higher accuracy using the voice of the subject.
[0072] (3) Third Embodiment Next, a third embodiment of the authentication device, the authentication method, and the recording medium will be described. Hereinafter, the third embodiment of the authentication device, the authentication method, and the recording medium will be described using the authentication system SYS to which the third embodiment of the authentication device, the authentication method, and the recording medium is applied. In the following description, the authentication system SYSa in the third embodiment is referred to as the authentication system SYSa to distinguish it from the authentication system SYS in the second embodiment.
[0073] Hereinafter, the authentication system SYSa in the third embodiment will be described with reference to FIG. 8. FIG. 8 is a block diagram showing the configuration of the authentication system SYSa in the third embodiment.
[0074] As shown in FIG. 8, the authentication system SYSa is different in that it includes a plurality of bone conduction microphones 2 as compared with the authentication system SYS. In the following description, an example in which the authentication system SYSa includes two bone conduction microphones 2 (specifically, bone conduction microphones 2#1 and 2#2) will be described as shown in FIG. 8. Other features of the authentication system SYSa may be the same as other features of the authentication system SYS.
[0075] The plurality of bone conduction microphones 2 are respectively arranged at a plurality of different positions with respect to the subject. For example, the bone conduction microphones 2 may be arranged so as to respectively contact a plurality of different parts of the subject. As an example, the bone conduction microphone 2#1 may be arranged so as to contact the head of the subject, and the bone conduction microphone 2#2 may be arranged so as to contact the ear of the subject or a part in the vicinity thereof. As an example of the bone conduction microphone 2#1 that contacts the head of the subject, there is a bone conduction microphone incorporated in a glasses-type wearable device (for example, the temple part of glasses). As an example of the bone conduction microphone 2#2 that contacts the ear of the subject or a part in the vicinity thereof, there is a bone conduction microphone incorporated in a headset-type wearable device that can be worn on the ear of the subject.
[0076] The use of one of the plurality of bone conduction microphones 2 and the use of another bone conduction microphone 2 different from the one of the plurality of bone conduction microphones 2 may be different. That is, the use of the bone conduction microphone 2#1 and the use of the bone conduction microphone 2#2 may be different. As an example, either one of the bone conduction microphones 2#1 and 2#2p may be used to calculate the registered feature amount registered in the collation DB 321. In this case, the registered feature amount may be calculated from the bone conduction sound detected by either one of the bone conduction microphones 2#1 and 2#2p. On the other hand, either the other of the bone conduction microphones 2#1 and 2#2p may be used to calculate the target feature amount for authenticating the subject. In this case, the calculation unit 311 provided in the authentication device 3 described above may calculate the target feature amount from the bone conduction sound detected by either the other of the bone conduction microphones 2#1 and 2#2p.
[0077] Here, the bone-conducted sound detected by the bone conduction microphone 2 may vary depending on the detection position of the bone-conducted sound. For example, the bone-conducted sound (particularly its feature amount) of a certain subject detected by the bone conduction microphone 2 disposed at one position may be different from the bone-conducted sound (particularly its feature amount) of the same subject detected by the bone conduction microphone 2 disposed at another position different from the one position. In this case, due to the difference between the bone conduction microphone 2 for calculating the registered feature amount and the bone conduction microphone 2 for calculating the target feature amount, the authentication accuracy of the authentication device 3 may deteriorate. Therefore, the authentication unit 312 provided in the above-described authentication device 3 may authenticate the subject in consideration of the difference in the position of the bone conduction microphone 2. Hereinafter, the authentication operation for authenticating the subject in consideration of the difference in the position of the bone conduction microphone 2 will be described with reference to FIG. 9. FIG. 9 is a flowchart showing the flow of the authentication operation for authenticating the subject in consideration of the difference in the position of the bone conduction microphone 2.
[0078] As shown in FIG. 9, also in the third embodiment, the calculation unit 311 acquires an air-conducted voice signal (step S11), the calculation unit 311 acquires a bone-conducted voice signal (step S12), and the calculation unit 311 calculates an air-conducted feature amount and a bone-conducted feature amount (step S13).
[0079] After that, the authentication unit 312 determines whether the position of the bone conduction microphone 2 with respect to the subject has changed as compared with the case where the registered feature amount has been calculated (step S31a). That is, the authentication unit 312 determines whether the position of the bone conduction microphone 2 used to calculate the registered feature amount is different from the position of the bone conduction microphone 2 used to calculate the target feature amount (that is, the position of the bone conduction microphone 2 when the operation shown in FIG. 9 is being performed, and the position of the bone conduction microphone 2 currently worn by the subject). To make this determination, in the collation DB 312, the registered feature amount may be associated with microphone position information regarding the position of the bone conduction microphone 2 used to calculate the registered feature amount. As a result, the authentication unit 312 can identify the position of the bone conduction microphone 2 used to calculate the registered feature amount by referring to the collation DB 312. Further, information regarding the position of the bone conduction microphone 2 used to calculate the target feature amount may be input to the authentication unit 312 by the subject, for example. Alternatively, the authentication unit 312 may estimate the position of the bone conduction microphone 2 currently worn by the subject (that is, the position of the bone conduction microphone 2 used to calculate the target feature amount) from the device number or the like of the bone conduction microphone 2 currently worn by the subject.
[0080] As a result of the determination in step S31a, if it is determined that the position of the bone conduction microphone 2 has changed (that is, the position of the bone conduction microphone 2 used to calculate the registered feature amount is different from the position of the bone conduction microphone 2 used to calculate the target feature amount) (step S31a: Yes), the authentication unit 312 corrects the bone conduction feature amount calculated in step S13 (step S32a). Specifically, the authentication unit 312 corrects the bone conduction feature amount so that the change in the bone conduction feature amount caused by the difference between the position of the bone conduction microphone 2 used to calculate the registered feature amount and the position of the bone conduction microphone 2 currently worn by the subject is offset. That is, the authentication unit 312 corrects the bone conduction feature amount so that the corrected bone conduction feature amount approaches (preferably coincides with) the bone conduction feature amount calculated when it is assumed that the position of the bone conduction microphone 2 currently worn by the subject is the same as the position of the bone conduction microphone 2 used to calculate the registered feature amount.
[0081] To correct the bone conduction feature amount, a correction parameter for correcting the bone conduction feature amount may be generated in advance from the difference between the feature amount of the bone conduction sound actually detected by the bone conduction microphone 2 arranged at one position and the feature amount of the bone conduction sound actually detected by the bone conduction microphone 2 arranged at another position different from the one position. For example, from the difference between the feature amount of the bone conduction sound actually detected by the bone conduction microphone 2#1 and the feature amount of the bone conduction sound actually detected by the bone conduction microphone 2#2, a correction parameter for correcting the feature amount of the bone conduction sound detected by the bone conduction microphone 2#1 to the feature amount of the bone conduction sound detected by the bone conduction microphone 2#2, and at least one of the correction parameters for correcting the feature amount of the bone conduction sound detected by the bone conduction microphone 2#2 to the feature amount of the bone conduction sound detected by the bone conduction microphone 2#1 may be generated in advance. In this case, the authentication unit 312 may correct the bone conduction feature amount using the correction parameter.
[0082] On the other hand, as a result of the determination in step S31a, if it is determined that the position of the bone conduction microphone 2 has not changed (that is, the position of the bone conduction microphone 2 used to calculate the registered feature amount is the same as the position of the bone conduction microphone 2 used to calculate the target feature amount) (step S31a: No), the authentication unit 312 may not correct the bone conduction feature amount calculated in step S13.
[0083] Thereafter, also in the third embodiment, the calculation unit 311 combines the air conduction feature amount calculated in step S13 and the bone conduction feature amount calculated in step S13 or corrected in step S32a (step S14), the calculation unit 311 calculates the target feature amount from the combined feature amount (step S15), and the authentication unit 312 authenticates the target person based on the target feature amount (step S16).
[0084] According to such a third embodiment, the authentication device 3 can appropriately authenticate the target person even when the position of the bone conduction microphone 2 used to calculate the registered feature amount is different from the position of the bone conduction microphone 2 used to calculate the target feature amount.
[0085] Still, FIG. 9 shows an authentication operation considering the difference in the position of the bone conduction microphone 2 in the first authentication operation described with reference to FIG. 4. However, even when the authentication device 3 performs the second authentication operation described with reference to FIG. 6, the authentication device 3 may consider the difference in the position of the bone conduction microphone 2. That is, even when the authentication device 3 performs the second authentication operation described with reference to FIG. 6, the authentication device 3 may correct the bone conduction feature amount in consideration of the difference in the position of the bone conduction microphone 2.
[0086] (4) Fourth Embodiment Next, a fourth embodiment of the authentication device, the authentication method, and the recording medium will be described. Hereinafter, the fourth embodiment of the authentication device, the authentication method, and the recording medium will be described using the authentication system SYS to which the fourth embodiment of the authentication device, the authentication method, and the recording medium is applied. Still, in the following description, the authentication system SYS in the third embodiment will be referred to as the authentication system SYSb to distinguish it from the authentication system SYS in the second embodiment.
[0087] The authentication system SYSb is different from the authentication system SYS in that a part of the second authentication operation is different. Other features of the authentication system SYSb may be the same as other features of the authentication system SYS.
[0088] Specifically, when performing the second authentication operation, the authentication device 3 authenticates the subject based on the air conduction feature amount (step S25 in FIG. 6), and authenticates the subject based on the differential feature amount (step S26 in FIG. 6). In the fourth embodiment, when the similarity between the air conduction feature amount and the first registered feature amount exceeds the authentication threshold (that is, it is determined that the subject matches the registered person), while the similarity between the differential feature amount and the second registered feature amount is below the authentication threshold (that is, it is determined that the subject does not match the registered person), it is presumed that some influence has occurred on the bone conduction feature amount. In this case, the authentication device 3 may correct the differential feature amount. For example, the bone conduction feature amount may vary depending on bone density. As an example, there may be a difference between the bone conduction feature amount of a person with normal bone density and the bone conduction feature amount of a person suffering from osteoporosis. In this case, when it is determined that the subject has osteoporosis, the authentication device 3 may correct the differential feature amount based on the information regarding the difference between the bone conduction feature amount of a person with normal bone density and the bone conduction feature amount of a person suffering from osteoporosis. As a result, even when some influence has occurred on the bone conduction feature amount, the authentication device 3 can appropriately authenticate the subject.
[0089] (5) Fifth Embodiment Subsequently, a fifth embodiment of the authentication device, the authentication method, and the recording medium will be described. Hereinafter, the fifth embodiment of the authentication device, the authentication method, and the recording medium will be described using the authentication system SYS to which the fifth embodiment of the authentication device, the authentication method, and the recording medium is applied. Note that in the following description, the authentication system SYS in the fifth embodiment is referred to as the authentication system SYSc to distinguish it from the authentication system SYS in the second embodiment.
[0090] The authentication system SYSc differs in that it may perform a weighting process on the bone conduction feature amount as compared with the authentication system SYS. Other features of the authentication system SYSc may be the same as other features of the authentication system SYS.
[0091] Specifically, the air conduction feature quantity is more susceptible to the influence of the ambient sound around the subject compared to the bone conduction feature quantity. Therefore, when the ambient sound around the subject is relatively large (for example, the magnitude of the ambient sound is greater than a threshold value), the weight of the bone conduction feature quantity may be increased compared to the case where it is not. Specifically, in the first authentication operation, the authentication device 3 may increase the weight of the bone conduction feature quantity when calculating the target feature quantity. In the second authentication operation, the authentication device 3 may increase the weight of the bone conduction feature quantity (in this case, actually, the weight of the bone conduction voice signal) when calculating the differential feature quantity. As a result, even when the ambient sound around the subject is relatively large, the authentication device 3 can appropriately authenticate the subject.
[0092] (6) Supplementary Note Regarding the embodiments described above, the following additional remarks are further disclosed. [Appendix 1] Calculating means for calculating an air conduction feature quantity, which is a feature quantity of the air conduction voice signal indicating the air conduction sound of the subject's voice, and a bone conduction feature quantity, which is a feature quantity of the bone conduction voice signal indicating the bone conduction sound of the subject's voice, from the air conduction voice signal and the bone conduction voice signal, and calculating a target feature quantity, which is a feature quantity of the subject's voice, by combining the air conduction feature quantity and the bone conduction feature quantity; Authentication means for authenticating the subject based on the target feature quantity An authentication device comprising the same. [Appendix 2] The calculating means calculates the target feature quantity using a neural network that outputs the target feature quantity when the combined air conduction feature quantity and bone conduction feature quantity are input. The authentication device according to Appendix 1. [Appendix 3] The calculating means calculates a differential feature quantity, which is a feature quantity of the difference between the frequency spectrum of the air conduction voice signal and the frequency spectrum of the bone conduction voice signal, The authentication means authenticates the subject based on the air conduction feature quantity and the differential feature quantity The authentication device according to Appendix 1 or 2. [Appendix 4] From an air-conducted voice signal indicating the air-conducted sound of the target person's voice and a bone-conducted voice signal indicating the bone-conducted sound of the target person's voice, an air-conducted feature quantity that is a feature quantity of the air-conducted voice signal, and a difference feature quantity that is a feature quantity of the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal are calculated, authentication means for authenticating the target person based on the air-conducted feature quantity and the difference feature quantity An authentication device comprising the same. [Appendix 5] The authentication means performs a first process of tentatively authenticating the target person based on the air-conducted feature quantity and a second process of tentatively authenticating the target person based on the difference feature quantity, and authenticates the target person definitively based on the result of the first process and the result of the second process The authentication device according to Appendix 4. [Appendix 6] From an air-conducted voice signal indicating the air-conducted sound of the target person's voice and a bone-conducted voice signal indicating the bone-conducted sound of the target person's voice, an air-conducted feature quantity that is a feature quantity of the air-conducted voice signal and a bone-conducted feature quantity that is a feature quantity of the bone-conducted voice signal are calculated, By combining the air-conducted feature quantity and the bone-conducted feature quantity, a target feature quantity that is a feature quantity of the target person's voice is calculated, The target person is authenticated based on the target feature quantity An authentication method. [Appendix 7] From an air-conducted voice signal indicating the air-conducted sound of the target person's voice and a bone-conducted voice signal indicating the bone-conducted sound of the target person's voice, an air-conducted feature quantity that is a feature quantity of the air-conducted voice signal and a difference feature quantity that is a feature quantity of the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal are calculated, The target person is authenticated based on the air-conducted feature quantity and the difference feature quantity An authentication method. [Appendix 8] On a computer, From an air-conducted voice signal indicating the air-conducted sound of the voice of a subject and a bone-conducted voice signal indicating the bone-conducted sound of the voice of the subject, an air-conducted feature amount that is a feature amount of the air-conducted voice signal and a bone-conducted feature amount that is a feature amount of the bone-conducted voice signal are calculated. By combining the air-conducted feature amount and the bone-conducted feature amount, a target feature amount that is a feature amount of the voice of the subject is calculated. Authenticate the subject based on the target feature amount. A recording medium on which a computer program for executing an authentication method is recorded. [Appendix 9] In a computer, From an air-conducted voice signal indicating the air-conducted sound of the voice of a subject and a bone-conducted voice signal indicating the bone-conducted sound of the voice of the subject, an air-conducted feature amount that is a feature amount of the air-conducted voice signal and a differential feature amount that is a feature amount of the difference between the frequency spectrum of the air-conducted voice signal and the frequency spectrum of the bone-conducted voice signal are calculated. Authenticate the subject based on the air-conducted feature amount and the differential feature amount. A recording medium on which a computer program for executing an authentication method is recorded.
[0093] At least a part of the constituent elements of each of the above-described embodiments can be appropriately combined with at least another part of the constituent elements of each of the above-described embodiments. It is not necessary to use some of the constituent elements of each of the above-described embodiments. Also, to the extent permitted by law, the disclosures of all documents (e.g., published gazettes) cited in this disclosure are incorporated by reference to form part of the description of this disclosure.
[0094] This disclosure can be appropriately modified within a range not contrary to the technical idea that can be read from the claims and the entire specification. An authentication device, an authentication method, and a recording medium accompanied by such modifications are also included in the technical idea of this disclosure.
Explanation of Reference Numerals
[0095] SYS Authentication System 1 Air-Conducted Microphone 2 Bone-Conducted Microphone 3. 1000 Authentication Device 31 Arithmetic Unit 311. 1001 Calculation Unit 312. 1002 Authentication Unit 32 Memory Device 321 Matching DB
Claims
1. Calculation means for calculating an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a bone conduction feature amount, which is a feature amount of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of a subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject, and combining the air conduction feature amount and the bone conduction feature amount so as to include at least one of one or more vector elements included in the air conduction feature amount and at least one of one or more vector elements included in the bone conduction feature amount, thereby calculating a target feature amount, which is a type of feature amount of the voice of the subject; authentication means for calculating a similarity between a registered feature amount that is registered and the target feature amount, and authenticating the subject based on the similarity; An authentication device comprising the above.
2. The calculation means calculates the target feature amount using a neural network that outputs the target feature amount when the combined air conduction feature amount and bone conduction feature amount are input. The authentication device according to Claim 1.
3. The calculation means calculates a difference feature amount, which is a feature amount of the difference between the frequency spectrum of the air conduction voice signal and the frequency spectrum of the bone conduction voice signal. The authentication means authenticates the subject based on the air conduction feature amount and the difference feature amount. The authentication device according to Claim 1 or 2.
4. The authentication means performs a first process of tentatively authenticating the subject based on the air conduction feature amount and a second process of tentatively authenticating the subject based on the difference feature amount, and finally authenticates the subject based on the result of the first process and the result of the second process. The authentication device according to Claim 3.
5. Calculating an air conduction feature amount, which is a feature amount of the air conduction voice signal, and a bone conduction feature amount, which is a feature amount of the bone conduction voice signal, from an air conduction voice signal indicating the air conduction sound of the voice of a subject and a bone conduction voice signal indicating the bone conduction sound of the voice of the subject; combining the air conduction feature amount and the bone conduction feature amount so as to include at least one of one or more vector elements included in the air conduction feature amount and at least one of one or more vector elements included in the bone conduction feature amount, thereby calculating a target feature amount, which is a type of feature amount of the voice of the subject; calculating a similarity between a registered feature amount that is registered and the target feature amount, and authenticating the subject based on the similarity; An authentication method.
6. On a computer, From an air-conducted voice signal indicating the air-conducted sound of the voice of the subject and a bone-conducted voice signal indicating the bone-conducted sound of the voice of the subject, an air-conducted feature amount that is a feature amount of the air-conducted voice signal and a bone-conducted feature amount that is a feature amount of the bone-conducted voice signal are calculated. By combining the air-conducted feature amount and the bone-conducted feature amount so as to include at least one of one or more vector elements included in the air-conducted feature amount and at least one of one or more vector elements included in the bone-conducted feature amount, a target feature amount that is a type of feature amount of the voice of the subject is calculated. The similarity between the registered registered feature amount and the target feature amount is calculated, and the subject is authenticated based on the similarity. A computer program for executing the authentication method.
Citation Information
Patent Citations
Device and method for estimating air-conducted sound
JP2004279768A
Personal identification system
JP2006010809A
Individual authentication system
JP2006011591A
Speech authentication device
JP2007017840A
Voice authentification system
JP2020184032A