An environmental noise-resistant gesture interaction action-independent identity authentication method

By constructing a personalized audio mapping model for the internal and external microphones of smart headphones and an adversarial learning network, facial tissue features are extracted, solving the reliability problem of teeth gesture interaction authentication under noise interference, realizing robust identity authentication independent of teeth gestures, and improving the accuracy and security of user identity identification.

CN121187528BActive Publication Date: 2026-06-26BEIJING UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF TECH
Filing Date
2025-09-05
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing tooth gesture interaction authentication methods are ineffective under environmental noise interference, and inconsistencies in tooth gestures affect authentication reliability and make it difficult to effectively resist imitation attacks.

Method used

By constructing a personalized audio signal mapping model for the internal and external microphones of smart headphones, facial tissue features are extracted using bone conduction pathways, and adversarial learning networks are combined to perform tooth and gesture-independent identity authentication, eliminating environmental noise interference and improving authentication robustness.

Benefits of technology

It achieves robust and reliable implicit authentication against inconsistencies in dental gestures under environmental noise, reduces the risk of impersonation attacks, and improves the accuracy and security of user identity verification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121187528B_ABST
    Figure CN121187528B_ABST
Patent Text Reader

Abstract

The anti-environmental noise tooth gesture interaction action independent identity authentication method belongs to the field of Internet of Things and ubiquitous computing. The ear outside and ear inside microphones of the intelligent earphone respectively perceive environmental noise signals and ear canal audio signals. The ear inside collected audio signals have frequency selection and individual difference characteristics due to the effect of ear closing plug. When the user performs tooth gesture interaction action, the corresponding ear inside bone conduction audio signals exist individual differences after modulation through the facial transmission path, and are not easily affected by the interaction action. The present application effectively eliminates environmental noise interference by constructing individualized mapping relationship of ear inside and outside audio signals, and uses the adversarial learning network to extract the transmission path tissue related biological features from the captured tooth gesture ear inside bone conduction audio signals, realizes the anti-environmental noise tooth gesture interaction action independent identity authentication. The present application has the advantages of fast authentication, effective resistance to environmental noise and robustness to inconsistent tooth gesture interaction action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of IoT wearable smart sensing technology, and relates to a method for dental interaction action-independent identity authentication that is resistant to environmental noise. Background Technology

[0002] With the continuous development of embedded technology, the types of sensors integrated into smart headphones are becoming increasingly diverse. For example, commercial smart headphones generally integrate motion sensors, in-ear microphones, and out-of-ear microphones. In addition, due to their convenient wearing and low cost, smart headphones have become widely popular and have broken through the traditional application scope of audio playback and recording, becoming a mobile intelligent computing platform that supports complex functions such as human-computer interaction, identity authentication, position estimation, and physiological signal monitoring.

[0003] During the use of various mobile applications based on smart headphones, these headphones inevitably access users' private information (such as location and physiological signals). Therefore, existing research has proposed numerous ear-worn user authentication methods for smart headphones to ensure user privacy and security. For example, EarGate captures bone conduction audio signals generated by a user's walking movements to extract gait biometrics for user authentication. MandiblePass and VoiceInEar capture bone conduction vibration and audio signals accompanying a user's speech to extract biometrics related to the user's head tissues for user authentication. However, these methods are only applicable to specific scenarios such as walking or voice communication, thus limiting their application scenarios. Furthermore, EarEcho and LR-Auth utilize passive perception of external environmental noise or active perception of audio playback (broadcast or music) in the ear to extract the geometric features of the ear canal shape for user authentication. While these methods do not restrict users to specific behaviors, they suffer from context-dependent authentication limitations. Meanwhile, how to provide implicit authentication synchronously for user interaction scenarios based on smart headphones has become a new research hotspot. For example, EarSlide and BudsAuth extract biometric features related to the user's biological tissues to achieve identity authentication during finger swipe gestures around the ear. However, finger swipe gestures in the ear area are easily intercepted by attackers in the environment, making them vulnerable to imitation attacks. Compared to facial finger swipe interactions, interaction methods based on teeth biting movements in the mouth are more covert. Therefore, ToothSonic and TeethPass proposed implicit identity authentication for teeth gesture interaction scenarios. First, ToothSonic uses an in-ear microphone to extract bone conduction audio signals caused by teeth gesture interactions, and achieves authentication by identifying the biometric features of tooth structure and conduction path tissues. However, this authentication method mainly relies on the isolation effect of the earphone shell against external noise to resist environmental noise interference, which has limited effectiveness in resisting environmental noise interference. In contrast, TeethPass uses an external microphone to capture environmental noise, and then eliminates the environmental noise collected by the in-ear microphone by modeling the conversion relationship between in-ear and out-of-ear noise. However, the in-ear and out-of-ear noise conversion relationship constructed by this method ignores the influence of frequency selection and individual differences. Furthermore, inconsistencies in users' dental gesture interactions can also affect the reliability of the aforementioned authentication methods. Therefore, there is an urgent need to design an implicit authentication method that is effective against external noise interference and robust to inconsistent dental gesture interactions. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing ear-worn identity authentication methods by constructing a personalized conversion model of audio signals between the user's inner and outer ear based on the external and internal microphones integrated into smart headphones. This model extracts user biometric features unrelated to the user's dental gesture interaction, thereby achieving a reliable implicit identity authentication method based on dental gesture interaction that resists environmental noise interference and requires no additional hardware.

[0005] The innovation of this invention lies in two aspects: First, due to the in-ear occlusion effect, the conversion mapping relationship between the user's in-ear and out-of-ear audio signals exhibits frequency selection and individual differences. Therefore, this invention designs an excitation signal based on a linear frequency modulated (LFM) signal to construct a personalized conversion mapping model for the user's in-ear and out-of-ear audio signals, effectively eliminating environmental noise interference. Second, the audio signals of teeth gestures collected by the in-ear microphone, after being conducted through the user's facial bone conduction, are not only related to the type of teeth gesture but also closely related to the facial tissues (e.g., bone and muscle density) involved in the bone conduction path. Specifically, the bone conduction audio signals of teeth gestures are modulated by the user's facial tissues during facial bone conduction (i.e., the signal is absorbed, reflected, and dispersed). Individual differences in the user's facial tissues are relatively stable during teeth gesture interactions and are not easily affected by the interaction. Therefore, this invention utilizes the modulation effect of the bone conduction path on the audio signals of teeth gestures to extract relevant biometric features of the user's facial tissues, achieving reliable implicit identity authentication independent of teeth gesture actions.

[0006] The objective of this invention is achieved through the following technical solutions.

[0007] Step 1: Signal Sample Collection. Sample collection specifically includes the collection of teeth gesture samples and the collection of preset excitation signal samples. First, during the registration and authentication phases, after the user wears the smart earphone, the authentication method of this invention utilizes its integrated external and internal ear microphones to simultaneously collect audio samples of the user's teeth gestures. Second, during the user registration phase, the user needs to use the external and internal ear microphones to collect audio samples for constructing the audio mapping relationship between the inside and outside of the ear.

[0008] Step 2: Construction of Intra- and Extra-ear Audio Mapping. During the registration phase, this step utilizes samples from chirp signals with different amplitude presets collected in Step 1 to construct the intra- and extra-ear audio mapping relationship. First, this step uses Fast Fourier Transform to process the audio samples collected by the in-ear and extra-ear microphones respectively, generating corresponding frequency domain amplitude spectra. Then, the mapping relationship between different pairs of in-ear and extra-ear samples at specific frequencies is modeled as a ratio relationship. Finally, after removing outliers, the average ratio of the frequency domain amplitude spectra of different samples is used as the final intra- and extra-ear audio mapping relationship.

[0009] Step 3: Environmental Noise Detection and Removal. Environmental noise interference may exist during the collection of audio samples of teeth gestures in the user registration and authentication phases. Therefore, this step first detects the presence of environmental noise based on the audio signals collected by the external ear microphone in the collected samples, and performs environmental noise removal if interference is present. First, this step uses the audio signal mean square decibel full-scale threshold detection method to determine the presence of environmental noise; if the value exceeds a preset threshold, environmental noise is considered to be present. Then, this step uses the user-personalized in-ear and out-of-ear audio mapping relationship constructed in Step 2 to convert the audio recorded by the external ear microphone into the audio recorded by the in-ear microphone. Finally, spectral subtraction is used for noise removal.

[0010] Step 4: Tooth gesture signal segmentation. This step uses a short-time energy-based method to segment the noise-reduced audio samples recorded by the in-ear microphone to obtain tooth gesture signal segments.

[0011] Step 5: Training the Tooth Gesture-Independent Authentication Model. During the registration phase, this step utilizes the user registration sample signal fragments obtained in Step 4 to train a tooth gesture-independent authentication model based on an adversarial learning network. This adversarial learning network mainly consists of three modules: a feature extractor, a user identity recognizer, and a tooth gesture recognizer. During the adversarial learning process, these three main modules are trained simultaneously. Afterwards, the user-personalized feature extractor and identity recognizer are saved (e.g., stored in the user authentication model database) for use in user identity verification during the authentication phase.

[0012] Step 6: User Identity Verification. During the authentication phase, the sample provided by the user to be authenticated undergoes environmental noise removal (Step 3) and signal segmentation (Step 4) to obtain a corresponding teeth gesture audio clip. This audio clip is then processed by the user's personalized feature extractor stored during the registration phase (Step 5) to extract user identity features. Finally, the user identity recognizer stored during the registration phase (Step 5) verifies whether the user to be authenticated is a registered user.

[0013] Beneficial effects:

[0014] Compared with the prior art, the present invention has the following advantages:

[0015] 1. This invention utilizes only the external and internal microphones commonly found in smart headphones to achieve reliable implicit authentication that effectively resists external environmental noise and is robust to inconsistent dental gesture interactions. This invention can work directly with existing dental gesture interaction methods, enabling reliable implicit authentication to be performed simultaneously during user dental gesture interactions.

[0016] 2. This invention further considers the characteristics of frequency selection and individual differences in the mapping relationship between ambient noise outside and inside the ear when the user is wearing smart headphones. By constructing a personalized audio mapping relationship model between the user's ear and outside the ear, environmental noise interference can be effectively eliminated.

[0017] 3. The present invention further achieves robust identity authentication that is independent of the user's tooth gesture interaction and robust to the problem of inconsistency in tooth gesture interaction by extracting the biometric features related to the facial tissue conduction path during the tooth gesture interaction process. Attached Figure Description

[0018] Figure 1 This is a diagram illustrating the architecture of the identity authentication method according to an embodiment of the present invention.

[0019] Figure 2 This is a prototype device for a smart earphone according to an embodiment of the present invention.

[0020] Figure 3 Adversarial learning authentication network as an embodiment of the present invention

[0021] Figure 4 The signal distortion ratio before and after denoising in this embodiment of the invention.

[0022] Figure 5 The average accuracy of the embodiments of the present invention

[0023] Figure 6 The false rejection rate in this embodiment of the invention

[0024] Figure 7 Error acceptance rate in embodiments of the present invention Detailed Implementation

[0025] like Figure 1 As shown, a noise-resistant tooth gesture interaction-independent authentication method includes the following steps:

[0026] Step 1: Signal Sample Collection

[0027] The sample collection process includes two sub-steps: tooth gesture sample collection and preset excitation signal sample collection.

[0028] Step 1.1: Collection of teeth gesture samples.

[0029] During the registration and authentication phases, after wearing the smart earphones, user u uses the integrated external and internal microphones to simultaneously collect audio samples of the user's teeth gestures.

[0030] in, and These represent the k-th audio sample collected by user u using the in-ear microphone and the external microphone when performing the g-th teeth gesture, respectively. The audio sampling rate is set to 48000Hz.

[0031] Step 1.2: Preset excitation signal sample collection.

[0032] While the shielding effect of the earphone shell can isolate ambient noise to some extent after a user wears smart earphones, residual ambient noise inside the ear still severely interferes with the audio signal of teeth gestures. Therefore, the authentication method of this invention utilizes external and internal ear microphones to eliminate ambient noise interference in the audio signal of teeth gestures collected during the registration and authentication phases. Specifically, during the user registration phase, the user uses the external and internal ear microphones to collect an audio sample set to construct an audio mapping relationship between the inside and outside of the ear. Specifically, while wearing smart earphones, the user places a smartphone at a certain distance (e.g., 10cm) outside the ear and actively plays audio at different amplitudes (A). a The preset Chirp audio signal (i.e., Chirp (A)) under the preset Chirp audio signal (i.e., Chirp (A)) a The system acquires audio samples from both the in-ear and out-of-ear microphones (audio sampling rate set to 48000Hz). Assuming M audio samples are collected for each amplitude value, the collected audio samples are denoted as follows: in and These represent the a-th amplitude (A) a The m-th audio signal sample recorded by the in-ear microphone and the out-of-ear microphone respectively under the preset Chirp audio signal. The preset Chirp audio signal (i.e., Chirp(A)) is the m-th audio signal sample recorded by the in-ear microphone and the out-of-ear microphone respectively. a Defined as:

[0033]

[0034] Where A a The amplitude is ∈{0.2, 0.4, 0.6, 0.8}. min =100Hz is the frequency modulation start frequency. max =2000Hz is the frequency modulation cutoff frequency. T = 2 seconds is the signal duration. t is a time variable and t∈[0,T]. This is the initial phase.

[0035] Step 2: Constructing Inner and Outer Ear Audio Maps

[0036] During the registration phase, this step utilizes preset Chirp signal samples corresponding to different amplitudes collected in step 1.2. Constructing the mapping relationship between in-ear and out-of-ear audio samples involves two sub-steps: generating the frequency domain amplitude spectrum of in-ear and out-of-ear audio samples and generating the mapping relationship between in-ear and out-of-ear audio samples.

[0037] Step 2.1: Generate the frequency domain amplitude spectrum of the inner and outer ear samples.

[0038] This step uses Fast Fourier Transform (FFT) to respectively... and The process is performed to generate the corresponding frequency domain amplitude spectrum. and

[0039]

[0040] Where |FFT(·)| represents taking the modulus of the output of FFT(·), i.e., calculating the frequency domain amplitude. Since the interaural noise caused by user body movements is generally distributed below 100Hz, and the frequency domain range of the audio signal for teeth gestures is mainly concentrated within 2000Hz, therefore... and Only retain the portion corresponding to the range f∈[100,2000].

[0041] Step 2.2: Generate the mapping relationship between the inner and outer ear audio frequencies.

[0042] This step first models the mapping relationship between the in-ear and out-of-ear microphone sample pairs at a specific frequency f as a ratio relationship. Then, after removing outliers using the median absolute deviation method (MAD(·)), the average of the contrast values ​​of different samples is used as the final in-ear and out-of-ear audio mapping relationship. That is:

[0043]

[0044] Where |A a | indicates the number of different amplitudes in the preset Chirp audio signal, and M indicates the number of audio samples collected for each amplitude.

[0045] Step 3: Environmental noise detection and elimination

[0046] Environmental noise may interfere with the collection of audio samples of teeth gestures during the user registration and authentication phases. Therefore, this step first relies on samples collected by an external ear microphone. Detect the presence of environmental noise. If noise interference is detected, environmental noise cancellation is necessary.

[0047] Step 3.1: Environmental noise detection.

[0048] Environmental noise in daily life (such as music, voice, or traffic and appliance noise) typically lasts for a relatively long time. Therefore, this invention utilizes a full-scale mean square decibel (dBFS) measurement based on audio signals. RMS The threshold detection method is used to determine the presence of environmental noise, i.e., when dBFS... RMSWhen the noise level exceeds a threshold θ, environmental noise is considered to exist, requiring subsequent environmental noise cancellation. Specifically,

[0049]

[0050] in, This indicates an audio sample recorded by an external microphone. The nth audio sample point, where N is The number of sample points. This invention empirically sets the threshold θ for determining the presence of environmental noise to -65 dBFS.

[0051] Step 3.2: Environmental noise elimination.

[0052] This step utilizes the user-personalized in-ear and out-of-ear audio mapping relationship constructed in step 2. Audio recorded by an external ear microphone The audio signal is converted to the audio recorded by the in-ear microphone. Then, noise cancellation is performed using spectral subtraction. Assume the in-ear audio signal containing external noise is... The signal contains residual external noise transmitted through the headphones, as well as audio from teeth gestures. It is first analyzed with a length of 1024 and an inter-frame repetition rate of 0.75. and Perform frame segmentation (label the i-th signal frame as follows) and Then, spectral subtraction denoising is performed frame by frame. Finally, the denoised signal frames are... Recombined into a pure signal with external noise eliminated The frame-by-frame denoising uses spectral subtraction as follows:

[0053]

[0054] In this context, FFT(·) and IFFT(·) represent Fast Fourier Transform and Inverse Fast Fourier Transform, respectively. |FFT(·)| represents taking the modulus of the output of FFT(·), i.e., calculating the frequency domain magnitude. phase(FFT(·)) represents taking the phase spectrum of the output of FFT(·). exp(·) represents the natural exponential function. j is the imaginary unit.

[0055] Step 4: Segmentation of teeth gesture signals:

[0056] This step uses a short-time energy-based method to record audio samples from the noise-cancelled in-ear microphone. The signal is segmented to obtain tooth gesture signal fragments. This mainly includes two sub-steps: short-time energy calculation and signal fragment segmentation.

[0057] Step 4.1: Short-time energy calculation.

[0058] Different types of dental gestures have significantly different durations. For example, lateral friction between teeth has a longer duration. Therefore, this step empirically selected 125ms as the frame duration and 0.75 as the inter-frame repetition rate for the samples. Perform frame segmentation. Assume each frame contains L sample points, and label the i-th frame as Fra. i [l], where l is the index of the intra-frame sample point and l = 1, 2, ..., L. Its short-time energy is...

[0059] Step 4.2: Signal segmentation.

[0060] The signal frame corresponding to the teeth gesture event has high short-time energy. Therefore, this step sets the threshold γ = 0.1 × max(STE(Fra)). i This is used to determine which signal frames contain teeth gesture events. Then, the first signal frame is matched with the last signal frame. The mid-sample points are used as the start and end points of the teeth gesture event, and are segmented to obtain the corresponding audio fragments. u .

[0061] Step 5: Training the teeth-gesture-independent authentication model:

[0062] During the registration phase, this step utilizes the audio fragment of the registered user's teeth gesture obtained in step 4. u An adversarial learning network is used to train a tooth gesture-independent authentication model. For example... Figure 2 As shown, the tooth gesture-independent authentication model training network used in this invention mainly includes three modules: a feature extractor, a user identity recognizer, and a tooth gesture recognizer. During adversarial learning, these three main modules are trained simultaneously. Afterwards, the user-personalized feature extractor and identity recognizer are saved (e.g., stored in a user authentication model database) for user identity verification during the authentication phase. The specific design is as follows:

[0063] Step 5.1: Feature Extractor.

[0064] This step first uses a short-time Fourier transform to generate an audio clip of the user's teeth gesture. uThe time-frequency spectrum of the input registration signal segment is used as input to the feature extractor. The feature extractor consists of three convolutional blocks and two pooling layers. The convolutional blocks transform the time-frequency spectrum of the input registration signal segment into a compressed feature representation through convolution operations. The pooling layers (max pooling) further reduce the dimensionality of the feature representation. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation layer. The output channels of the convolutional layers in the three convolutional modules are 16, 32, and 64, respectively, and the kernel size is 3×3 for each. The pooling layer window is set to 2×2. The feature representations output by the feature extractor will be input into the subsequent user identification and tooth gesture recognition systems, respectively.

[0065] Step 5.2: User Identifier.

[0066] The user identity identifier, used to improve user authentication capabilities, consists of two convolutional blocks, two fully connected layers, and one...

[0067] The system consists of Sigmod activation layers. Each convolutional block comprises a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max-pooling layer. The user identification function trains a binary classifier to determine whether an input sample is a registered or unregistered user. The two convolutional modules have 64 and 32 output channels respectively, and both have a 3×3 kernel size. The Sigmod activation layer outputs the probability that the input sample is a registered user. Y r This is the true label for whether the sample is a registered user. If the sample is a registered user, Y... r The value is 1; if the user is not a registered user, Y... r The value is 0. Therefore, the loss function of the user identification detector is:

[0068]

[0069] Step 5.3: Tooth gesture recognition device.

[0070] A teeth gesture recognizer is used to reduce the ability of extracted features to represent teeth gesture types during adversarial learning. This module consists of a convolutional block, fully connected layers, and softmax activation layers. The convolutional block includes convolutional layers, batch normalization layers, ReLU activation layers, and max-pooling layers. The convolutional layers have 32 output channels, and the kernel size is 3×3. The max-pooling layer window is set to 1×1. Its loss function is defined using cross-entropy.

[0071]

[0072] in, The label for the i-th tooth gesture is 1 (1 indicates that it belongs to the tooth gesture, and 0 indicates that it does not).

[0073] Let G be the predicted probability of the tooth gesture recognizer for i tooth gestures. |G| is the number of types of tooth gestures involved.

[0074] Step 5.4 Adversarial Network Training

[0075] In adversarial learning, the feature extractor, user identity recognizer, and teeth gesture recognizer are trained jointly. Specifically, adversarial learning training achieves teeth gesture-independent feature extraction by minimizing the gesture recognition accuracy. Furthermore, adversarial learning training improves user identity feature extraction capability by maximizing the accuracy of the user identity recognizer. Therefore, given the teeth gesture recognizer loss function (L... g ) and the user identity recognition loss function (L r After that, the overall loss function for adversarial learning is:

[0076] L Final =αL r -βL g #(9)

[0077] Here, α and β are empirically set to 0.8 and 0.2, respectively.

[0078] Step 6: User Identity Verification

[0079] During the authentication phase, the sample provided by the user to be authenticated undergoes environmental noise removal (step 3) and signal segmentation (step 4) to obtain a corresponding teeth gesture audio segment. This audio segment is then processed by the user-personalized feature extractor stored in the registration phase (step 5) to extract user identity features. Finally, the user identity recognizer stored in the registration phase (step 5) determines whether the user to be authenticated is a registered user. Specifically, in a single-user authentication scenario, the registered user probability is output by the user identity recognizer (binary classifier) ​​stored in the registration phase. If the user to be authenticated is identified as a registered user, then the user is considered an unregistered user. In multi-user authentication scenarios, the registration phase generates and stores a personalized feature extractor and a corresponding user identity identifier for each registered user. If multiple identity identifiers for registered users exist... In this case, the probability of identifying the user to be authenticated as the output registered user is the highest. Registered users.

[0080] Example verification:

[0081] Because manufacturers of commercial smart headphones do not provide interfaces to acquire audio data recorded by the external and internal ear microphones. Therefore, as Figure 3As shown, this invention utilizes a commercial lavalier microphone and earphone shell to assemble a prototype device supporting in-ear and out-of-ear audio acquisition, in order to test and evaluate the authentication method involved in this invention. A total of 20 volunteers were recruited to complete sample collection in the laboratory. Four volunteers were randomly assigned to play the role of attackers, while the remaining 15 volunteers played the role of registered users. In the process of generating a personalized teeth-gesture-independent authentication model for each registered user, the constructed training set included 20 samples from the registered user (as positive samples) and 20 samples from each of the other 15 users (as negative samples) to train the user identity recognition device. The constructed test samples included the remaining samples from the registered user (playing the role of a registered user during authentication) and all samples from the four attackers (playing the role of strangers to launch random attacks). Furthermore, to evaluate environmental noise resistance, the experiment used speakers to play popular music to simulate environmental noise and construct a test dataset. Specifically, two volunteers were selected, and using their respective in-ear and out-of-ear audio mapping relationships, they denoised the simulated environmental noise to obtain residual in-ear noise. A test dataset of "real out-of-ear noise - real in-ear noise - in-ear noise generated through mapping relationships" was constructed.

[0082] Regarding the elimination of ambient noise in the ear: This invention utilizes the signal distortion ratio (i.e., the ratio of the pure ambient tooth gesture audio to the residual noise signal energy after denoising) to quantitatively evaluate the denoising effect. For example... Figure 4 As shown, without denoising, a large amount of residual environmental noise remains in the ear, resulting in a relatively low signal distortion ratio, with a median of only -5.31 dB. If noise cancellation is performed directly using external noise as a reference (subtracting the audio recorded by the external microphone from the audio recorded by the internal microphone), the median signal distortion ratio is -4.05 dB, indicating that the denoising effect is not ideal. If the actual internal ear noise is used as a reference (eliminating the actual internal ear noise from the audio recorded by the internal microphone), the median signal distortion ratio is significantly improved (reaching 14.00 dB). However, using the external ear noise processing based on the mapping relationship between the internal and external ear audio involved in this invention, and using the generated internal ear noise as a reference, the median signal distortion ratio is 13.59 dB, which is very close to the denoising effect using the actual internal ear noise as a reference. Therefore, the experimental results demonstrate that the environmental noise cancellation technology involved in this invention can effectively suppress internal ear environmental noise.

[0083] Regarding authentication accuracy: the average accuracy rate (i.e., the average of the accuracy rate for registered user authentication and the accuracy rate for stranger authentication) is used to evaluate the authentication accuracy of the above authentication methods. For example... Figure 5As shown, after eliminating environmental noise interference, the median average accuracy rate for all registered users reached 97.42%. Furthermore, even the lowest average accuracy rate was above 90%, indicating that the authentication method of this invention can accurately distinguish registered users from strangers. In addition, without noise cancellation, the average accuracy rate distribution was more dispersed, with the median decreasing to 64.89%, further demonstrating that the environmental noise cancellation method applied in this invention can effectively eliminate environmental noise.

[0084] Regarding the false rejection rate: The false rejection rate reflects the user experience of registered users. It is defined as the proportion of registered users who are identified as strangers. For example... Figure 6 As shown, before environmental noise cancellation, the median false rejection rate reached 70.22%, indicating that environmental noise seriously interfered with the authentication of registered users. In contrast, after environmental noise cancellation, the median false rejection rate decreased to 5.17%, demonstrating that the environmental noise cancellation method involved in this patent can effectively ensure the user experience of registered users.

[0085] Regarding the false acceptance rate: The false acceptance rate directly reflects the security of the authentication system and is defined as the proportion of samples where a stranger is identified as an authenticated user. For example... Figure 7 As shown, the median false acceptance rate is 0% both before and after denoising. Furthermore, the distribution of the false acceptance rate after denoising is clustered at a lower level. This demonstrates that the authentication method of the present invention can effectively resist random attacks and protect user privacy and security.

Claims

1. A noise-resistant, tooth gesture interaction-independent authentication method, characterized in that: Step 1: Signal sample collection; Sample collection specifically includes the collection of teeth gesture samples and the collection of preset excitation signal samples; First, during the registration and authentication phases, after the user wears the smart earphone, the integrated external and internal microphones simultaneously collect audio samples of the user's teeth gestures. Second, during the user registration phase, the user needs to use the external and internal microphones to collect audio samples for constructing the audio mapping relationship between the inside and outside of the ear. Step 2: Construction of Intra- and Extra-ear Audio Mapping; During the registration phase, samples under different preset Chirp signals collected in Step 1 are used to construct the intra- and extra-ear audio mapping relationship. First, this step uses Fast Fourier Transform to process the audio samples collected by the intra- and extra-ear microphones respectively to generate the corresponding frequency domain amplitude spectra. Then, the mapping relationship between different pairs of intra- and extra-ear samples at a specific frequency is modeled as a ratio relationship. Finally, after removing outliers, the average value of the frequency domain amplitude spectrum ratios of different samples is used as the final intra- and extra-ear audio mapping relationship. Step 3: Environmental Noise Detection and Removal; First, based on the audio signals collected by the external ear microphone in the collected samples, the presence of environmental noise is detected, and environmental noise removal is performed if noise interference is present. First, this step uses the audio signal mean square decibel full-scale threshold detection method to determine the presence of environmental noise; that is, if the value exceeds a preset threshold, environmental noise is considered to be present. Then, using the user-personalized in-ear and out-of-ear audio mapping relationship constructed in Step 2, the audio recorded by the external ear microphone is converted to the audio recorded by the in-ear microphone. Finally, noise removal is performed using spectral subtraction. Step 4: Tooth gesture signal segmentation; The audio samples recorded by the noise-canceling in-ear microphone are segmented using a short-time energy-based method to obtain tooth gesture signal segments; Step 5: Training the teeth-gesture-independent authentication model; During the registration phase, the user registration sample signal fragments obtained in step 4 are used to train the tooth gesture-independent authentication model based on an adversarial learning network. This adversarial learning network consists of three modules: a feature extractor, a user identity recognizer, and a teeth gesture recognizer. During the adversarial learning process, the above three modules are trained simultaneously; afterwards, the user-personalized feature extractor and identity recognizer are saved for user identity verification during the authentication phase. Step 6: User authentication; During the authentication phase, the sample provided by the user to be authenticated undergoes environmental noise cancellation and signal segmentation to obtain the corresponding tooth gesture audio segment; this audio segment is then processed by the user's personalized feature extractor stored during the registration phase to extract the user's identity features. Then, the user identity identifier stored during the registration phase verifies whether the user to be authenticated is a registered user.

2. The method according to claim 1, characterized in that: Step 1: Signal Sample Collection Sample collection specifically includes two sub-steps: tooth gesture sample collection and preset excitation signal sample collection; Step 1.1: Collection of teeth gesture samples; During the registration and authentication phases, users When wearing smart earphones, the integrated external and internal microphones simultaneously collect audio samples of the user's teeth gestures. ;in, and Representing users respectively Using in-ear and external microphones to perform the first The first tooth was collected under the hand gesture 1 audio sample; audio sampling rate set to 48000. ; Step 1.2: Preset excitation signal sample collection; During the user registration phase, the user Audio sample sets are collected using external and internal ear microphones to construct an audio mapping relationship between the inside and outside of the ear; specifically, when wearing smart headphones, the user places a smartphone at a certain distance outside the ear and actively plays audio at different amplitudes. The preset Chirp audio signal is Audio samples were acquired from the in-ear and external microphones; a total of [number] samples were collected for each amplitude. If there are 10 audio samples, then the collected audio samples are denoted as 10. ;in and They represent the first individual amplitude Under the corresponding preset Chirp audio signal, the in-ear microphone and the external microphone recorded the first... One audio signal sample; the preset Chirp audio signal is... Defined as: ; in The amplitude; This is the starting frequency for frequency modulation. This is the frequency modulation cutoff frequency; Seconds represent the duration of the signal; It is a time variable and =0 represents the initial phase.

3. The method according to claim 2, characterized in that: Step 2: Constructing Inner and Outer Ear Audio Maps During the registration phase, preset Chirp signal samples corresponding to different amplitudes collected in step 1.2 are used. Constructing the mapping relationship between in-ear and out-of-ear audio; specifically including two sub-steps: generating the frequency domain amplitude spectrum of in-ear and out-of-ear audio samples and generating the mapping relationship between in-ear and out-of-ear audio. Step 2.1: Generate the frequency domain amplitude spectrum of the inner and outer ear samples; Using Fast Fourier Transform To each and The process is performed to generate the corresponding frequency domain amplitude spectrum. and : ; ; in, Indicates to The output result is moduloed, i.e., the frequency domain amplitude is calculated; the interaural noise caused by the user's limb movements is generally distributed in... The following, and the frequency range of the audio signal for teeth gestures is mainly concentrated in Within; therefore, and Only keep The part corresponding to the range; Step 2.2: Generation of audio mapping relationship between inside and outside the ear; This step first involves setting a specific frequency. The mapping relationship between sample pairs collected by the in-ear and out-of-ear microphones was modeled as a ratio relationship; then, the median absolute deviation method was used. After removing outliers, the average of the contrast values ​​of different samples is used as the final in-ear and out-of-ear audio mapping relationship; that is: ; in This indicates the number of different amplitude values ​​in the preset Chirp audio signal. This indicates the number of audio samples collected for each amplitude value.

4. The method according to claim 3, characterized in that: Step 3: Environmental noise detection and elimination Based on samples collected by external ear microphones Detect the presence of environmental noise; if noise interference is detected, environmental noise elimination is required. Step 3.1: Environmental noise detection; Using full-scale mean square decibels of audio signals The threshold detection method is used to determine the presence of environmental noise, that is, when At that time, it was determined that environmental noise existed and subsequent environmental noise reduction work was necessary; specifically... ; in, This indicates an audio sample recorded by an external microphone. The Middle One audio sample point, for Number of sample points; threshold Set to -65dBFS; Step 3.2: Environmental noise elimination; This step utilizes the user-personalized in-ear and out-of-ear audio mapping relationship constructed in step 2. Audio recorded by an external ear microphone The audio signal is converted into audio recorded by the in-ear microphone; then, noise cancellation is performed using spectral subtraction; assuming the in-ear audio signal containing external noise is... The signal contains residual external noise transmitted through the headphones, as well as audio of teeth gestures; it is first processed with a length of 1024 and an inter-frame repetition rate of 0.

75. and Perform frame processing and mark the first The signal frames are respectively and Then, spectral subtraction denoising is performed frame by frame; finally, the denoised signal frames are... Recombined into a pure signal with external noise eliminated The frame-by-frame denoising uses spectral subtraction as follows: ; in, and These represent the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively. Indicates to The modulus of the output result is calculated, i.e., the frequency domain amplitude is calculated. Indicates to Calculate the phase spectrum from the output results; Represents the natural exponential function; It is the imaginary unit.

5. The method according to claim 4, characterized in that: Step 4: Segmentation of teeth gesture signals: This step uses a short-time energy-based method to record audio samples from the noise-cancelled in-ear microphone. Segmentation is performed to obtain segments of tooth gesture signals; It includes two sub-steps: short-time energy calculation and signal segmentation; Step 4.1: Short-time energy calculation; We selected 125ms as the frame duration and 0.75 as the inter-frame repetition rate for the samples. Perform frame segmentation; assume each frame includes The nth sample point is labeled with the nth sample point. Each frame is For intra-frame sample point index and Its short-term energy is ; Step 4.2: Signal segmentation; The signal frame corresponding to the teeth gesture event has high short-time energy; therefore, this step sets a threshold. This is used to determine which signal frames contain the teeth gesture event; then, the first signal frame is matched with the last signal frame. The mid-sample points were used as the start and end points of the teeth gesture event, and the event was segmented to obtain the corresponding audio segments. .

6. The method according to claim 5, characterized in that: Step 5: Training the teeth-gesture-independent authentication model: During the registration phase, the audio clip of the registered user's teeth gesture obtained in step 4 is used. A tooth gesture-independent authentication model is trained based on an adversarial learning network. The training network comprises three modules: a feature extractor, a user identity recognizer, and a tooth gesture recognizer. During the adversarial learning process, these three main modules are trained simultaneously. Subsequently, the user-personalized feature extractor and identity recognizer are saved for user identity verification during the authentication phase. The specific design is as follows: Step 5.1: Feature Extractor; This step first uses short-time Fourier transform to generate an audio clip of the user's teeth gesture. The time-frequency spectrum of the input registered signal segment is used as input to the feature extractor. The feature extractor consists of three convolutional blocks and two pooling layers. The convolutional blocks transform the time-frequency spectrum of the input registered signal segment into a compressed feature representation through convolution operations. The pooling layers further reduce the dimensionality of the feature representation. Each convolutional block consists of a convolutional layer, a batch normalization layer, and a ReLU activation layer. The number of output channels of the convolutional layers in the three convolutional modules are 16, 32, and 64, respectively, and the spatial size of the convolutional kernel is 16, 32, and 64, respectively. The pooling layer window is set to The feature representations output by the feature extractor will be input into the subsequent user identity recognizer and tooth gesture recognizer, respectively. Step 5.2: User Identifier; The user identification unit, designed to enhance user authentication capabilities, consists of two convolutional blocks, two fully connected layers, and one Sigmad activation layer. Each convolutional block comprises a convolutional layer, a batch normalization layer, a ReLU activation layer, and a max-pooling layer. The user identification unit trains a binary classifier to determine whether an input sample is a registered or unregistered user. The two convolutional modules have 64 and 32 output channels respectively, and both have convolutional kernels with a spatial size of [missing information]. The Sigmoid activation layer output is the probability that the input sample is a registered user. This is the true label for whether the sample is a registered user; if the sample is a registered user, The value is 1; if the user is not a registered user, The value is 0; therefore, the loss function of the user identification detector is: ; Step 5.3: Tooth gesture recognition device; A tooth gesture recognizer is used to reduce the ability of extracted features to represent tooth gesture types during adversarial learning. This module consists of a convolutional block, fully connected layers, and a Softmax activation layer. The convolutional block includes convolutional layers, batch normalization layers, ReLU activation layers, and max pooling layers. The convolutional layers have 32 output channels, and the convolutional kernel space size is [missing information]. The maximum pooling layer window is set to... Its loss function is defined using cross-entropy: ; in, For the first A label for each tooth gesture; 1 indicates that the gesture belongs to that tooth gesture, and 0 indicates that it does not. For teeth gesture recognition The predicted probability of a tooth gesture; This refers to the number of different types of dental gestures. Step 5.4 Training the Adversarial Network Given a tooth gesture recognizer loss function With user identity recognition loss function Then, the overall loss function for adversarial learning is: ; in, and Set to 0.8 and 0.

2.

7. The method according to claim 6, characterized in that: Step 6: User Identity Verification During the authentication phase, the sample provided by the user to be authenticated undergoes environmental noise cancellation and signal segmentation to obtain the corresponding tooth gesture audio segment; this audio segment is then processed by the user's personalized feature extractor stored in the registration phase to extract the user's identity features; and then, the user identity recognizer stored in the registration phase determines whether the user to be authenticated is a registered user. In a single-user authentication scenario, the probability of a registered user is stored during the registration phase, output by the user identity recognizer. If the user is not registered, the system will identify them as a registered user; otherwise, they will be considered an unregistered user. In multi-user authentication scenarios, the registration phase generates and stores a corresponding personalized feature extractor and a corresponding user identity recognizer for each registered user; if multiple registered user identity recognizers exist... In this case, the probability of identifying the user to be authenticated as the output registered user is the highest. Registered users.

Citation Information

Patent Citations

  • User identity authentication method using inward microphone of earphone

    CN115348049A