Personal authentication method, personal authentication device, and personal authentication program
The personal authentication system uses multiple sensors to capture speech and body vibrations, integrating time-frequency features and calculating differences to enhance authentication accuracy in noisy environments and account for user variability, addressing the limitations of single-method sound collection.
Patent Information
- Application Number
- JP2021137909
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing personal authentication methods using speech recognition face challenges in noisy environments and are influenced by the physical condition of the speaking user, relying on a single sound collection method, which affects authentication performance.
A personal authentication system utilizing multiple sensors, including microphones and vibration sensors, to capture speech and body vibrations at different locations, integrating time-frequency features and calculating differences to enhance authentication accuracy.
Improves authentication performance by stabilizing identification through diverse data collection and analysis, reducing errors and enhancing user satisfaction.
Smart Images

Figure 0007814723000002 
Figure 0007814723000003 
Figure 0007814723000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a personal authentication method, a personal authentication device, and a personal authentication program for performing a personal authentication process using a user's speech. [Background technology]
[0002] In recent years, speech recognition has come to be used in a variety of situations, such as sorting goods in logistics sites, inputting addresses for deliveries, and inputting medical records in medical settings. Accordingly, there is a growing need for personal authentication using user speech to check whether a person holds a license at work sites or to identify individuals to prevent information leaks.
[0003] Generally, personal authentication using speech involves collecting voice with a microphone, performing acoustic analysis on the input voice to calculate features, and comparing the features with pre-registered personal patterns to determine which individual the input voice resembles, thereby authenticating the individual, as in Patent Document 1. However, authentication performance deteriorates in noisy environments, so bone conduction microphones or headsets with built-in bone conduction microphones are used to record clear voices.
[0004] Patent document 2 performs personal authentication by extracting personal features from signals received by a bone conduction microphone through the body (skeleton) of the person to be authenticated by a feature extraction unit that performs frequency analysis of the bone conduction sound, and then comparing the personal features with personal data registered in a feature database by a feature matching unit.
[0005] The biometric authentication device described in Patent Document 3 transmits a signal pattern to a living body, receives a response signal that is transmitted through biological tissue, calculates the transmission characteristics of the signal transmitted through the living body based on the signal pattern transmitted to the living body and the response signal received from the living body, extracts a feature amount that is a quantity that differs for each living body based on the calculated transmission characteristics, and identifies an individual by comparing the extracted feature amount with pre-stored feature amounts. [Prior art documents] [Patent documents]
[0006] [Patent Document 1] Patent Publication No. 2021-64110 [Patent Document 2] Japanese Patent Application Laid-Open No. 2003-58190 [Patent Document 3] Japanese Patent Application Laid-Open No. 2008-40882 Summary of the Invention [Problem to be solved by the invention]
[0007] The problem in the above-mentioned conventional examples is how to accurately grasp the individuality of a user from his / her voice in order to improve the performance of personal identification, and various ideas are described in the respective patent documents.
[0008] That is, in the prior art described in Patent Document 1, a user speaks a predetermined string of characters, the voice is picked up through a microphone, this input voice is subjected to acoustic analysis to calculate features, and these are compared with pre-registered personal patterns to determine which individual the input voice resembles, thereby authenticating the individual.
[0009] However, this type of personal authentication method uses the individuality inherent in a person's voice, i.e., individuality, to perform personal authentication, but it simply utilizes the difference between the input voice and a pre-registered voice, and depending on the speaking user's physical condition, the voice may differ significantly from the registered voice, making it difficult to achieve stable authentication performance.In addition, this personal authentication method only uses one type of sound collection method: voice recorded with a microphone.
[0010] On the other hand, in Patent Document 2, a bone conduction microphone is used as the microphone, thereby preventing the user's voice from being buried in surrounding noise and significantly reducing authentication performance.
[0011] However, even if a bone conduction microphone is used, it is still the same in that personal authentication is essentially performed using the individuality inherent in the voice, i.e., the individuality, and it may be difficult to achieve stable identification performance depending on the physical condition of the speaking user, etc. Furthermore, this personal authentication method only uses one type of sound collection method, that is, the sound recorded by the bone conduction microphone.
[0012] Furthermore, the biometric authentication device of Patent Document 3 transmits a signal to a living body, receives a response signal that is transmitted through the living body tissue, calculates the transfer characteristics of the living body based on the signal transmitted to the living body and the response signal received from the living body, and performs individual identification by comparing the feature amount extracted for each living body based on the transfer characteristics with the feature amount stored in advance.
[0013] However, this biometric authentication method uses a method in which sound output from a speaker is input into the body, and then the sound that passes through the body is recorded by a microphone, but only one type of acoustic data is used. Sound passing through the body is affected by various factors such as physique, constitution, and body composition. Therefore, the transmission characteristics of sound as it passes through the body also vary, and may change depending on the sound path.
[0014] The present invention addresses this diversity in personal authentication and aims to improve authentication performance based on speech. [Means for solving the problem]
[0015] The present inventors have conducted extensive research to solve the above problems and have found that the following inventions meet the above objectives, thereby completing the present invention.
[0016] <1> a plurality of sensors for sensing vibrations caused by the speaker's speech; a feature extraction unit that extracts time-frequency feature values of each of the plurality of sensor data acquired by the plurality of sensors; a feature integrated data creation unit that creates feature integrated data that combines the time-frequency feature values extracted from each of the sensor data; and an authentication unit that authenticates the speaker using the feature integrated data and a registered model including the feature integrated data that has been registered in advance in a database. <2> the feature integrated data creation unit includes an inter-feature difference data calculation unit that calculates inter-feature difference data that is a difference between the time-frequency feature values extracted from each of the sensor data; <1> The personal authentication device described in <3> the plurality of sensors are different types of sensors, The sensor is two or more types of sensors selected from the group consisting of a microphone, a bone conduction microphone, a vibration sensor, and an acceleration sensor. <1> or <2> The personal authentication device described in <4> The plurality of sensors acquires information at different positions, The location is targeted at two or more locations selected from the group consisting of the mouth, chin, throat, nape of the neck, around the ears, and ear canal. <1> ~ <3> 10. The personal authentication device according to claim 9, wherein <5> The feature amount difference data calculation unit is configured to calculate a difference and / or a ratio between the feature amount difference data. <2> The personal authentication device described in <6> the authentication unit performs authentication based on a statistical distance between the feature integrated data of the speaker to be authenticated and the feature integrated data registered in the registered model and / or a similarity compared with a trained model by machine learning. <1> ~ <5> The personal authentication device described in <7> the feature integrated data creation unit includes a spectrum analysis means for performing a spectrum analysis for each unit time on the sounds acquired at the installation positions of the plurality of sensors; The difference or ratio of the spectral analysis results of the mounting positions of the plurality of sensors obtained by performing the spectral analysis by the spectral analysis means. <1> ~ <6> 10. The personal authentication device according to claim 9, wherein <8> an acquiring step of acquiring vibrations based on the speaker's speech using a plurality of sensors; a feature extraction step of extracting a plurality of time-frequency features from each of the plurality of sensor data acquired by the plurality of sensors; a feature integrated data creation step of creating feature integrated data by aggregating time-frequency feature values of each of the sensor data acquired by the plurality of sensors; and an authentication step of performing authentication using the feature integrated data and a registered model that has been registered in a database in advance. <9> a feature extraction unit that extracts a plurality of time-frequency features from a plurality of sensor data acquired by a plurality of sensors that sense vibrations based on the speech of a speaker; a feature integrated data creation unit that creates feature integrated data by aggregating time-frequency feature values of each of the sensor data acquired by the plurality of sensors; A program for causing a computer to function as an authentication unit that performs authentication using the feature integrated data and a registered model that has been registered in a database in advance. [Effects of the Invention]
[0017] The present invention can significantly improve authentication performance by processing a plurality of pieces of speech-based data together, which can also increase the satisfaction of users of personal authentication devices. [Brief explanation of the drawings]
[0018] [Figure 1] FIG. 2 is a flow diagram of an example of an authentication method of the present invention. [Figure 2] FIG. 1 is a processing block diagram of the first embodiment. [Figure 3] FIG. 10 is a processing block diagram of a second embodiment. [Figure 4] FIG. 11 is a processing block diagram of a third embodiment. [Figure 5] FIG. 10 is a diagram illustrating an example of creating feature difference data during authentication. [Figure 6]FIG. 10 is a diagram illustrating an example of creating feature difference data during learning. [Figure 7] FIG. 10 is a diagram illustrating an outline of comparison during authentication. [Figure 8] FIG. 10 is a diagram illustrating an example of mounting positions of a plurality of sensors. [Figure 9] FIG. 2 is a diagram showing the attachment position of the sensor in the ear canal. [Figure 10] FIG. 2 is a diagram illustrating an example of the structure of a sensor. [Figure 11] FIG. 10 is a diagram illustrating an example of analysis according to an embodiment. [Figure 12] FIG. 10 is a diagram illustrating an example of analysis according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0019] The following describes in detail an embodiment of the present invention, but the following description of the constituent elements is one example (typical example) of an embodiment of the present invention, and the present invention is not limited to the following content unless the gist of the present invention is changed. Note that when the expression "to" is used in this specification, it is used as an expression that includes the numerical values before and after it.
[0020] [Personal authentication device of the present invention] The personal authentication device of the present invention includes a plurality of sensors that sense vibrations accompanying the speech of a speaker; a feature extraction unit that extracts time-frequency features from each of a plurality of pieces of sensor data acquired by the plurality of sensors; a feature integrated data creation unit that creates integrated feature data that compiles the time-frequency features extracted from each of the sensor data; and an authentication unit that authenticates the speaker using the integrated feature data and a registered model that includes the integrated feature data and that has been registered in a database in advance.
[0021] [Personal authentication method of the present invention] The personal authentication method of the present invention includes an acquisition step of acquiring data using a plurality of sensors that sense vibrations based on the speaker's speech; a feature extraction step of extracting a plurality of time-frequency features from each of the plurality of sensor data acquired by the plurality of sensors; a feature-integrated data creation step of creating integrated feature data by aggregating the time-frequency features of each of the sensor data acquired by the plurality of sensors; and an authentication step of performing authentication using the integrated feature data and a registered model that has been registered in a database in advance.
[0022] [Personal authentication program of the present invention] The personal authentication program of the present invention is a program for causing a computer to function as a feature extraction unit that extracts multiple time-frequency features from each of multiple sensor data acquired by multiple sensors that sense vibrations based on the speech of a human speaker, a feature integrated data creation unit that compiles the time-frequency features of each of the sensor data acquired by the multiple sensors to create integrated feature data, and an authentication unit that performs authentication using the integrated feature data and a registered model that has been registered in a database in advance.
[0023] In addition, in this application, the personal authentication method of the present invention can be performed using the personal authentication device of the present invention, and the personal authentication program of the present invention can also be used for these, and the corresponding configurations in this application can be used mutually.
[0024] FIG. 1 is a flow diagram of an example of an authentication method according to the present invention. The authentication method according to the present invention can be implemented by a personal authentication device according to an embodiment described later. First, in step S11, human speech is recorded using multiple sensors that sense speech at various locations. Next, in step S21, multiple feature extraction is performed to extract multiple time-frequency features from the recorded sensor data. In step S31, feature integration data is created. In step S41, authentication is performed using a registered model. In step S51, the authentication result is notified.
[0025] [Get] The authentication of the present invention utilizes sensor data related to voice acquired by multiple sensors. This sensor data can be acquired by multiple sensors that sense the voice uttered by the speaker. These sensors are placed in multiple different positions or multiple different types of sensors are combined to sense voice based on human speech and vibrations transmitted within the body, such as bone conduction, at various locations.
[0026] [Feature Extraction] In the authentication process of this invention, time-frequency features are extracted from the sensor data recorded by multiple sensors. Time-frequency features are features related to the relationship between time and frequency of vibrations emitted by the speaker and detected by the sensors.
[0027] [Feature integration data creation] For authentication in this invention, feature integrated data is created by summarizing the time-frequency features extracted from each sensor data. This is to facilitate authentication by integrating data acquired from sensors at multiple locations and types to obtain information that is difficult to obtain from a single sensor. The integrated data can be obtained by adding, subtracting, multiplying, dividing, or combining these for each feature.
[0028] [certification] In the authentication of the present invention, speaker authentication is performed using feature integrated data and a registered model including the feature integrated data that has been registered in advance in a database. Information to be authenticated is registered in advance as a registered model. Authentication is performed by comparing the information with the registered model to determine similarity, etc., and identifying the information as exceeding a predetermined threshold. The authentication result can then be notified or displayed. The authentication result can also be used as a signal to start another operation.
[0029] An example of application of the present invention is the construction industry. In recent years, the Japanese construction industry has been working to improve the working environment at construction sites and other locations due to a decrease in the number of workers. For example, the use of professional headsets equipped with bone conduction microphones that enable clear voice calls even in noisy environments has been expanding. Therefore, the present invention aims to improve the performance of speaker identification using audio sensors such as bone conduction microphones. Furthermore, speaker identification can also be applied to methods for verifying license possession at work sites and for identifying individuals to prevent information leaks. The present invention can also be realized as application software that runs on personal computers, tablet devices, smartphones, etc. based on sensed vibrations.
[0030] [Embodiment 1] 2 shows a processing block diagram of embodiment 1. Personal authentication device 31 is a device that first senses the voice and vibrations emitted by the person using the personal authentication device with first sensor 11 and second sensor 12, and then performs processing for authentication.
[0031] The personal authentication device 31 has a plurality of sensors. The plurality of sensors in the personal authentication device 31 include a first sensor 11 and a second sensor 12. The first sensor 11 and the second sensor 12 can be various microphones that record human speech, such as a condenser microphone, a bone conduction microphone, or a throat microphone. It is also possible to use a sensor that captures vibrations and movements generated by the human body, such as a vibration sensor or an acceleration sensor.
[0032] Although this embodiment shows an example in which two sensors, a first sensor and a second sensor, are used, the number of sensors may be greater, such as three or more, or four or more. By using different types of sensors or arranging the sensors in different positions, speech with different time-frequency features can be acquired and used even for the same speech content by the same speaker.
[0033] By recording bone conduction sounds, which are less susceptible to noise, at multiple locations and calculating feature difference data using the multiple recorded bone conduction sounds, it is possible to reduce errors in the feature difference data that occur with each recording and obtain stable authentication performance. For this reason, it is preferable that the sensor is capable of sensing bone conduction sounds.
[0034] The first feature extraction unit 13 extracts time-frequency features from the sound or vibration data captured by the first sensor 11. The second feature extraction unit 14 extracts time-frequency features from the sound or vibration data captured by the second sensor 12. Here, the time-frequency features may be various features that express the time-frequency characteristics of the signal captured by the sensor, such as a power spectrum, logarithmic power spectrum, or mel-logarithmic power spectrum calculated by frequency analysis of the sound or vibration, prediction coefficients obtained by a parametric method such as linear predictive analysis, or cepstrum coefficients obtained by homomorphic analysis.
[0035] The feature difference data calculation unit 15 is an example of a feature integrated data creation unit. The feature difference data calculation unit 15 calculates, as feature difference data, the difference between the time-frequency feature extracted by the first feature extraction unit 13 and the time-frequency feature extracted by the second feature extraction unit 14. Examples of the difference from the time-frequency feature include a method of calculating by division when a power spectrum is used as the feature, and a method of calculating by subtraction when a logarithmic power spectrum is used as the feature. In other words, the feature difference data calculation unit 15 can be configured to calculate, as feature difference data, the difference or difference between the time-frequency feature extracted from each of the sensor data.
[0036] In addition to these, any of various arithmetic operations such as addition, subtraction, multiplication, and division, or a combination of these, may be used as integrated data. In particular, if the difference or ratio of features from each sensor is used as feature difference data, it becomes easier to extract the transfer characteristics of the living body contained in the time-frequency features, including those that acquire voice transmitted via a living body.
[0037] The feature difference data calculation unit 15 may include a spectral analysis means for performing spectral analysis on the sounds recorded at the attachment positions of the plurality of sensors at predetermined unit time intervals. The spectral analysis means may perform spectral analysis on the sounds recorded at the attachment positions of the plurality of sensors, and the difference or ratio of the spectral analysis results obtained by the spectral analysis may be used. This unit time interval may be, for example, 5 to 50 msec, or 10 to 25 msec.
[0038] The authentication difference data storage unit 16 stores the feature difference data calculated by the feature difference data calculation unit 15. At this time, for example, it is possible to calculate and store the average of the feature difference data over multiple time frames, or to calculate and store statistical data such as the mean and variance.
[0039] The authentication unit 17 performs authentication to identify which individual the feature difference data stored in the authentication-time difference data storage unit 16 belongs to. There are various conceivable methods for the processing by the authentication unit 17, including, for example, a method using a statistical distance measure such as a discriminant function, Euclidean distance, or Mahalanobis distance, and a method using a machine learning model such as a decision tree, SVM, or neural network.
[0040] The enrollment model unit 18 stores an enrollment model for each individual to be used when performing authentication in the authentication unit 17. The enrollment model to be stored in the enrollment model unit 18 corresponds to the authentication method performed by the authentication unit 17, and is created in advance by learning using feature difference data of the individual to be authenticated. Data for this enrollment model can be data that has been registered in advance on the same device, or data that has been registered in advance on a device with a common configuration. Typical speech samples recorded by speakers who are expected to be used as the authentication target are used for the enrollment model.
[0041] For example, if a discriminant function or statistical distance measure is used as the registered model, this corresponds to statistical quantities such as the discrimination threshold, mean, and variance. On the other hand, if a machine learning model such as a decision tree, SVM (support vector machine), or neural network is used as the registered model, this corresponds to thresholds, SVM parameters, neural network models, etc.
[0042] The authentication results can be communicated, used, or recorded by any means. For example, information identifying the speaker of the authentication result can be communicated using sound or an image. Alternatively, the authentication results can be used as a signal to start or stop the device or system to be operated.
[0043] [Embodiment 2] The second embodiment shows a more detailed example of an embodiment based on the first embodiment. In the second embodiment, the operation of the personal authentication device will be explained by dividing it into two major cases, namely, during authentication and during learning. The operation of the device during authentication refers to the operation when a person who wishes to undergo personal authentication actually performs personal authentication using the voice or body movements that he or she utters himself or herself. On the other hand, the operation of the device during learning refers to the operation when a registration model is created in advance by learning using feature difference data of the person who wishes to be authenticated. Therefore, the following explanation will be divided into the operation during authentication and the operation during learning.
[0044] FIG. 3 shows a processing block diagram during authentication in the second embodiment. In the personal authentication device 32, the authentication / learning changeover switch (state during authentication) 2 indicates that the device is operating during authentication. First, the first microphone 101 is a specific example of a microphone used as the first sensor 1 shown in the first embodiment, and the second microphone 102 is a specific example of the second sensor 2 shown in the first embodiment. As specific examples of these, various microphones can be used, such as a condenser microphone, a bone conduction microphone, or a throat microphone.
[0045] In this embodiment, amplifiers 103, 104, A / D converters 105, and 106 are used. In this embodiment, an example is shown in which the first and second microphones are used, but other microphones such as those described in the first embodiment may also be used.
[0046] The first feature extraction unit 107 extracts time-frequency features from the sound captured by the first microphone 101. Similarly, the second feature extraction unit 108 extracts time-frequency features from the sound captured by the second microphone 102. As in the first embodiment, the time-frequency features may be various features that express the time-frequency characteristics of the signal captured from the microphones, such as a power spectrum, logarithmic power spectrum, or Mel-logarithmic power spectrum calculated by frequency analysis of sound or vibration, prediction coefficients obtained by a parametric method such as linear predictive analysis, or cepstrum coefficients obtained by homomorphic analysis.
[0047] Next, the feature difference data calculation unit 109 calculates, as feature difference data, the difference between the time-frequency feature extracted by the first feature extraction unit 107 and the time-frequency feature extracted by the second feature extraction unit 108. As in the first embodiment, the difference from the time-frequency feature can be calculated by various methods of addition, subtraction, multiplication, and division, such as a division method when a power spectrum is used as the feature, or a subtraction method when a logarithmic power spectrum is used as the feature.
[0048] The authentication difference data storage unit 110 stores the feature difference data calculated by the feature difference data calculation unit 109. As in the first embodiment, for example, the feature difference data may be averaged over multiple time frames and stored, or statistical data such as the mean and variance may be calculated and stored.
[0049] The authentication unit 111 performs authentication to identify which individual the feature difference data stored in the authentication-time difference data storage unit 110 belongs to. Various methods can be considered for the processing of the authentication unit 111, and as in the first embodiment, various methods can be considered, such as a method using a statistical distance measure such as a discriminant function, Euclidean distance, or Mahalanobis distance, or a method using a machine learning model such as a decision tree, SVM, or neural network.
[0050] The enrollment model unit 114 stores an enrollment model for each individual to be used for authentication by the authentication unit 111. The enrollment model to be stored in the enrollment model unit 114 differs depending on the authentication method used by the authentication unit 111, but is created in advance by learning using feature difference data of the individual to be authenticated.
[0051] For example, when a discriminant function or a statistical distance measure is used as a registered model as in the first embodiment, statistical quantities such as a discrimination threshold, a mean, and a variance correspond to this. On the other hand, when a machine learning model such as a decision tree, an SVM, or a neural network is used as a registered model, thresholds, SVM parameters, a neural network model, etc. correspond to this.
[0052] Here, the learning-time difference data storage unit 112 and the learning unit 113 are processes used only during learning, and when the authentication / learning changeover switch (authentication state) 2 is in the authentication state, as shown by the dotted lines in Figure 3, there is no flow of information to the learning-time difference data storage unit 112, the learning unit 113, or further to the registered model unit 114, and no processing is performed by the learning-time difference data storage unit 112 or the learning unit 113. Note that learning processing may be performed while authentication is being performed, and simultaneous processing of authentication and learning is not prevented.
[0053] Figure 4 shows a processing block diagram during learning in embodiment 2. In personal authentication device 32, authentication / learning changeover switch (learning state) 2 indicates that the device is operating during learning. The operations in Figures 3 and 4 are similar except for the processes indicated by dotted lines, so only the differences between Figures 3 and 4 will be explained here.
[0054] First, the learning-time difference data storage unit 112 stores the feature difference data calculated by the feature difference data calculation unit 109. At this time, as in the first embodiment, it is possible to calculate and store the average of the feature difference data over multiple time frames, or to calculate and store statistical data such as the mean and variance. Furthermore, feature difference data for various times and situations is accumulated in advance.
[0055] The learning unit 113 performs a learning process to create a registration model required for authentication processing using the feature difference data stored in the learning difference data storage unit 112. There are various possible methods for the processing by the learning unit 111, but as in the first embodiment, when a statistical distance measure such as a discriminant function, Euclidean distance, or Mahalanobis distance is used as the authentication processing method, statistical quantities such as a discrimination threshold and a mean and variance are learned as the registration model.
[0056] On the other hand, when a machine learning model such as a decision tree, SVM, or neural network is used as the authentication processing method, thresholds, SVM parameters, neural network models, etc. are learned as registered models. In either case, the registered model unit 114 is created in the learning unit 113 from the learning difference data accumulated in the learning difference data storage unit 112.
[0057] Here, the authentication difference data storage unit 110 and the authentication unit 111 are processes used only during authentication, and when the authentication / learning changeover switch (learning state) 2 is in the learning state, as shown by the dotted lines in Fig. 4, no information flows to the authentication difference data storage unit 110 or the authentication unit 111, and no processing is performed by the authentication difference data storage unit 110 or the authentication unit 111. Note that the learning process may be performed while the authentication process is being performed.
[0058] [Embodiment 3] 5 shows a method for creating feature difference data during authentication according to embodiment 2. In personal authentication device 32, authentication / learning changeover switch (state during authentication) 2 indicates that the device is operating during authentication.
[0059] First, the first microphone 101 is a specific example of an ear canal microphone used as the microphone shown in embodiment 2, and various microphones can be used, such as a normal condenser microphone, a bone conduction microphone, or a throat microphone.
[0060] On the other hand, the second microphone 102 is a specific example of a case where a bone conduction microphone is used as the microphone shown in the second embodiment, and various microphones such as a normal condenser microphone or a throat microphone can be used.
[0061] Fig. 5(a) shows a specific example of a case where a logarithmic power spectrum is used as the time-frequency feature extracted by the first feature extracting unit 107 in the second embodiment. Similarly, Fig. 5(b) shows a specific example of a case where a logarithmic power spectrum is used as the time-frequency feature extracted by the second feature extracting unit 108.
[0062] In Fig. 5, the operation of subtracting Fig. 5(b) from Fig. 5(a) corresponds to a method of subtracting logarithmic power spectra obtained from two microphones as feature difference data. This feature difference data is stored in the authentication difference data storage unit 110. At this time, various methods are possible, such as a method of averaging the logarithmic power spectra obtained from each microphone over multiple time frames and then subtracting, or a method of subtracting the logarithmic power spectra obtained from each microphone and then averaging over multiple time frames.
[0063] FIG. 6 is a diagram showing a method for creating feature difference data during learning in embodiment 2. In personal authentication device 32, authentication / learning changeover switch (learning state) 2 indicates that the device is operating during learning. The operations in FIGS. 5 and 6 differ only in whether they are for authentication processing or learning processing, and therefore the processes in FIGS. 5 and 6 are identical and will not be described further.
[0064] Fig. 7 is a diagram showing an authentication method according to the second embodiment. Fig. 7(a) is a specific example of using a logarithmic power spectrum as a time-frequency feature when creating feature difference data according to the second embodiment. On the other hand, Fig. 7(b) is a specific example of using a logarithmic power spectrum as a time-frequency feature when creating the registered model unit 114 according to the second embodiment.
[0065] FIG. 7 shows a method for performing authentication in the authentication processing unit 111 using feature difference data calculated using the logarithmic power spectrum stored in the authentication difference data storage unit 110, and a registered model similarly trained using the logarithmic power spectrum as a time-frequency feature.
[0066] [Embodiment 4] Fig. 8 is a perspective view of a personal authentication device according to the second embodiment, Fig. 9 is a diagram showing the mounting positions of multiple microphones according to the second embodiment, and Fig. 10 is a diagram showing an example of the structure of a sensor according to the second embodiment. In the perspective view of the ear canal microphone according to the second embodiment, a body case 1 is connected to a wearing body 3 made of a soft material, and the tip of the wearing body 3 is provided with a first microphone 101 used as an ear canal microphone.
[0067] A personal authentication device 32 is built into the main body case 1. An authentication / learning changeover switch 2 and a second microphone 102 used as a bone conduction microphone are provided on the surface of the main body case 1. Furthermore, as shown in Figures 9 and 10, a first microphone 101 and a receiver 4 are integrated into a wearing body 3 made of a soft material, and the wearing body 3 and the first microphone 101 open into the ear canal 8 via an acoustic tube 5, and the receiver 4 opens into the ear canal 8 via an acoustic tube 6.
[0068] By wearing the main body case 1 along the back side of the ear, the second microphone 102 installed on the surface of the main body case 1 comes into contact with the cartilage on the back side of the ear.
[0069] In Figure 9, the main case 1 is not attached to the ear (pinna) 7 so that the positional relationship of the receiver 4 and the first microphone 101 relative to the ear canal 8 can be easily understood. However, in actual use, the main case 1 is hung over the ear (pinna) 7 and positioned along the back of the ear (pinna) 7, and the receiver 4 and the first microphone 101 are attached to the entrance of the ear canal 8 or inserted into this ear canal 8 as shown in Figure 9.
[0070] The authentication / learning selector switch 2 is operated by the user when registering (learning) his / her personal information. Note that the operation of this authentication / learning selector switch 2 may be automatically switched depending on the personal authentication device's detection of a human voice or the internal processing status. In this case, the authentication / learning selector switch 2 is not necessary. [Example]
[0071] The following experiments relating to the present invention were carried out.
[0072] An example of the flow of the speaker identification method proposed by this invention is explained below. First, the Mel logarithm power spectrum (logMel) of bone conduction sound and laryngeal sound is calculated, and the average value for all frames of the input speech is found. Next, the difference value of logMel between bone conduction sound and laryngeal sound is calculated, and whether or not the speaker is the relevant speaker is identified using an SVM that has been trained in advance using the difference value of a specific speaker.
[0073] The spectral difference between bone-conducted and pharyngeal sounds removes the common speech spectrum contained in each. The remaining spectrum may contain transfer characteristics due to the shape and tissue of the head and neck. We believe that these transfer characteristics may contain speaker-specific characteristics.
[0074] By collecting input voices from different parts of the body using multiple sound collection methods, such as bone conduction and pharyngeal sounds, various acoustic characteristics of human speech can be utilized, and authentication can be performed more efficiently. accuracy On the other hand, it is possible to use airway sounds other than bone conduction sounds and pharyngeal sounds as input speech, but airway sounds are easily affected by ambient noise and authentication is difficult. accuracy Therefore, we propose a speaker identification method that utilizes spectral differences using bone conduction and throat microphones. Below, we conduct an evaluation through identification experiments.
[0075] In this study, we propose a speaker identification method using the spectral difference between bone conduction (BC) and pharyngeal (TH) sounds. In the proposed method, the Mel log power spectrum of BC and TH is calculated. The spectral difference is then calculated and speaker identification is performed using SVM. The proposed method using the spectral difference between BC and TH achieved the highest identification rate, confirming the effectiveness of the spectral difference between BC and TH.
[0076] [Discrimination experiment overview] Acoustic features are calculated from speech waves using openSMILE. The spectral differences between bone-conducted and laryngeal sounds are calculated from the obtained acoustic features, and an SVM is trained to conduct a recognition experiment to determine whether or not the speaker is the correct one. Based on the recognition results, the effectiveness of the spectral differences between bone-conducted and laryngeal sounds is judged. For comparison, three types of recognition using each voice (AC, BC, TH) alone are also evaluated.
[0077] [Audio data] In this study, three types of sounds are used: air conduction sound (AC), bone conduction sound (BC), and pharyngeal sound (TH), and they are recorded simultaneously using the following three microphones. Condenser microphone (SONY ECM-530, unidirectional) Bone conduction microphone (Temco Japan EM21N-Tip) Throat microphone (Retivis 1 Pin 3.5mm Throat Mic Earpiece Covert Air Tube Earpiece for Phones)
[0078] Speech recordings were conducted in a soundproof room with six 21-year-old adult males. The recorded speech data consisted of 20 sentences with 503 ATR phoneme balance sentences. The speech was recorded digitally at a sampling frequency of 48 kHz and a quantization bit rate of 16 bits. To analyze the speech with openSMILE, the audio was downsampled to a sampling frequency of 16 kHz using ffmpeg.
[0079] [Experimental conditions] The average 0-7 dimension logMel of all frames of the input speech is extracted using openSMILE. The sampling frequency of the extracted speech is 16 kHz, the analysis frame length is 25 ms, and the shift width is 10 ms. The difference value calculated from the obtained 0-7 dimension logMel is used as a feature. In this study, this data is used to identify six speakers (m1-m6) using an SVM in Weka [Reference URL Weka: https: / / www.cs.waikato.ac.nz / ml / weka / ]. This is a pattern method that separates multiple classes by learning data. The following formula uses a polynomial kernel function, and the SVM settings are as shown in Table 2, where K(x, x´) is the kernel function and the others are constants.
[0080]
number
[0081] A confusion matrix was used to evaluate the experiment. Accuracy and F-measure were calculated from this. Accuracy is the proportion of correct answers in all data. F-measure is the harmonic mean of recall and accuracy. In addition, since the evaluation was performed using 10-fold cross-validation, 18 sentences were used for training data and 2 sentences for evaluation data for each speaker. A confusion matrix is used to evaluate the experiment. Accuracy and F-measure can be calculated from the confusion matrix. Accuracy is the percentage of correct answers among all data. F-measure is the harmonic mean of recall and precision. Also, because the number of data is very small, 10-fold cross-validation was performed, with each speaker having 18 sentences for training and 2 sentences for evaluation.
[0082] [Identification results and discussion] Identification results The average accuracy and F-measure values for each voice are shown in Figures 11 and 12. The classification results using the difference between BC and TH (98.3%) were higher than the classification results using AC alone (83.3%), BC alone (95.0%), and TH alone (94.1%). Furthermore, there was no significant difference in the classification results between the results using BC alone and the results using TH alone. However, the classification results using AC alone were much lower than the other classification results.
[0083] ·Consideration The discrimination result using the difference between BC and TH (98.3%) was higher than the discrimination results using AC alone (83.3%), BC alone (95.0%), and TH alone (94.1%). This confirmed the effectiveness of the speaker discrimination method using the spectral difference between bone conduction and pharyngeal sounds.
[0084] The average values of the features (0-7 dimension logMel) obtained from AC, BC, and TH for each speaker were evaluated. The difference between the average values of the BC and TH features (0-7 dimension logMel) was evaluated. It was found that the differences in the feature values between speakers were relatively small. On the other hand, it was found that the differences in the feature values between speakers were relatively large.
[0085] Furthermore, in the classification results, misrecognition was more likely when the difference in feature values between each speaker was small, and less likely when the difference in feature values between each speaker was large. Furthermore, BC and TH contain transfer characteristics because they recorded speech that had traveled through the human body. However, AC does not contain transfer characteristics because it recorded speech that had traveled through the air. Therefore, compared to BC and TH, AC has smaller differences in feature values between each speaker, making it less likely to be misrecognized, which is thought to have resulted in lower classification results.
[0086] As described above, we proposed a speaker identification method using the spectral difference between bone-conducted sound and laryngeal sound, and evaluated it through identification experiments. As a result, the identification result using the difference value between bone-conducted sound and laryngeal sound achieved 98.3%, confirming the effectiveness of the speaker identification method using the spectral difference between bone-conducted sound and laryngeal sound. We also found that speech containing transfer characteristics due to the shape and tissues of the head and neck (bone-conducted sound, laryngeal sound) had higher identification accuracy than speech not containing transfer characteristics (air-conducted sound). [Industrial Applicability]
[0087] The present invention can be used for personal authentication of speakers and is industrially useful. [Explanation of symbols]
[0088] 1 Main unit case 11 First Sensor 12 Second Sensor 13, 107 First feature extraction unit 14, 108 Second feature extraction unit 15, 109 Feature difference data calculation section 16, 110 Authentication difference data storage section 17, 111 Authentication Department 18, 114 Registered Model Section 101 First Microphone 102 Second microphone 103, 104 Amplifier 105, 106 A / D converter 112 Learning difference data storage unit 113 Learning Department 2 Authentication / Learning Switch 31, 32 Personal authentication device 3. Wearable body 4 receivers 5, 6 Acoustic tube 7. Ear (Auricle) 8 Ear canal
Claims
1. a plurality of sensors for sensing vibrations caused by the speaker's speech; a feature extraction unit that extracts time-frequency feature values of each of the plurality of sensor data acquired by the plurality of sensors; a feature integrated data creation unit that creates feature integrated data that combines the time-frequency feature values extracted from each of the sensor data; an authentication unit that authenticates the speaker by using the feature integrated data and a registered model including the feature integrated data that has been registered in advance in a database; A personal authentication device, wherein the sensor data includes at least sensor data of bone conduction sound and sensor data of pharyngeal sound.
2. 2. The personal authentication device according to claim 1, wherein the feature integrated data creation unit includes an inter-feature difference data calculation unit that calculates inter-feature difference data that is a difference between the time and frequency features extracted from each of the sensor data.
3. the plurality of sensors are different types of sensors, 3. The personal authentication device according to claim 1, wherein the sensor is two or more types of sensors selected from the group consisting of a microphone, a bone conduction microphone, a vibration sensor, and an acceleration sensor.
4. 3. The personal authentication device according to claim 2, wherein the feature difference data calculation unit calculates a difference and / or a ratio between the feature difference data.
5. The personal authentication device according to any one of claims 1 to 4, wherein the authentication unit performs authentication based on a statistical distance between the feature integrated data of the speaker to be authenticated and the feature integrated data registered in the registered model and / or a similarity between the feature integrated data and the registered model as compared with a trained model by machine learning.
6. the feature integrated data creation unit includes a spectrum analysis means for performing a spectrum analysis for each unit time on the sounds acquired at the installation positions of the plurality of sensors; 6. The personal authentication device according to claim 1, wherein the spectral analysis means performs spectral analysis to obtain a difference or ratio of spectral analysis results for the attachment positions of the plurality of sensors.
7. an acquiring step of acquiring vibrations based on the speaker's speech using a plurality of sensors; a feature extraction step of extracting a plurality of time-frequency feature values from each of the plurality of sensor data acquired by the plurality of sensors; a feature integrated data creation step of creating feature integrated data by aggregating time / frequency feature amounts of each of the sensor data acquired by the plurality of sensors; an authentication step of performing authentication using the feature amount integrated data and a registered model registered in a database in advance, A personal authentication method, wherein the sensor data includes at least sensor data of bone conduction sounds and sensor data of pharyngeal sounds.
8. a feature extraction unit that extracts a plurality of time-frequency feature values from a plurality of sensor data acquired by a plurality of sensors that sense vibrations based on the speech of a speaker; a feature integrated data creation unit that creates feature integrated data by aggregating time and frequency feature amounts of each of the sensor data acquired by the plurality of sensors; a program for causing a computer to function as an authentication unit that performs authentication using the feature amount integrated data and a registered model that has been registered in a database in advance, The sensor data includes at least sensor data of bone conduction sound and sensor data of pharyngeal sound.
Citation Information
Patent Citations
Communications equipment
JP2000102087A
Personal authentication system
JP2003058190A
Voice processing apparatus and voice processing method
JP2003264883A
Individual authentication system
JP2006011591A
Mobile terminal, communication system, voice muffling method, program and recording medium
JP2006166300A