A method for continuous authentication of a smartphone based on structural conducted acoustic signals
By sensing the structural acoustic signals of the user's hand posture through the built-in microphone and speaker of the smartphone, and combining them with a CNN model for identity authentication, the technology solves the problems of vulnerability to attack and limited application scenarios in existing technologies, and achieves fast, secure and continuous identity authentication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF TECH
- Filing Date
- 2025-08-20
- Publication Date
- 2026-04-21
AI Technical Summary
Existing smartphone authentication methods are vulnerable to attacks and require users to perform specific actions, making it difficult to meet continuous authentication needs and limiting their application scenarios.
By utilizing the built-in microphone and speaker of a smartphone, the system actively senses the structural acoustic signals related to the user's hand holding posture through audio perception, and combines this with a CNN model for identity authentication, extracting the biometric features of the user's hand holding posture.
It achieves fast, hardware-free, and widely applicable continuous identity authentication with high security and accuracy, effectively resisting imitation attacks.
Smart Images

Figure CN121078168B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a smartphone authentication method based on structurally conducted acoustic signals. Specifically, it relates to a user authentication method that utilizes the built-in microphone and speaker of a smartphone to obtain acoustic signals conducted by the internal structure of the phone through active audio perception and extract biometric features related to the user's hand holding posture. This method belongs to the fields of Internet of Things and ubiquitous technology. Background Technology
[0002] To prevent the leakage of personal privacy information, smartphones support various authentication methods. For example, authentication methods based on users' facial and fingerprint biometrics are widely used in commercial smartphones. However, these methods are vulnerable to identity theft attacks. Furthermore, while voiceprint-based authentication allows users to periodically update their voice command information, it is only applicable in limited scenarios (i.e., it cannot be used in places where quiet is required). In contrast, PIN code and pattern-based authentication methods are favored by users due to their speed and user-friendly experience. However, in recent years, a series of attacks targeting these authentication methods have emerged, seriously threatening their security. To further improve smartphone authentication security, various authentication methods based on novel biometrics have emerged in recent years. For example, research has utilized on-screen finger interaction gestures and teeth occlusion voiceprints to provide instant authentication. However, these methods require users to perform specific actions during the authentication process, reducing the user experience and making it difficult to meet continuous authentication needs. Therefore, there is an urgent need to design a continuous authentication method that is fast, widely applicable, and requires no additional hardware. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing smartphone authentication methods. Based on the microphones and speakers that are commonly integrated into smartphones, this invention proposes a continuous identity authentication method that is highly secure, fast, requires no additional hardware, and has a wide range of applications by sensing the biometric features related to the user's hand posture while holding the smartphone.
[0004] The innovation of this invention lies in the fact that, in the everyday scenario of a user holding a mobile phone with one hand, the user's hand-holding posture exhibits individual differences, which can be used as a biometric for identity authentication. On one hand, there are individual differences in the size and thickness of a user's palm; on the other hand, during the process of holding the phone, the contact position and area between the user's palm and the phone are also influenced by individual usage habits. These two aspects make the user's single-handed holding posture highly individualized, providing feasibility for its use as a biometric for identity authentication. Furthermore, the structural sound signals transmitted within the phone are affected by the user's palm contact position and area, exhibiting differences. Therefore, this invention utilizes the aforementioned effects related to the transmission of structural sound signals within the phone, extracting these sound signal features to characterize the user's hand-holding posture-related biometrics, thereby achieving secure user identity authentication.
[0005] The objective of this invention is achieved through the following technical solutions.
[0006] Step 1: Sensing signal generation.
[0007] This step involves designing an acoustic sensing signal consisting of a pilot section, a sensing section, and an idle section to capture acoustic conduction signals from the internal structure of the mobile phone.
[0008] Step 1.1: Design of the pilot section.
[0009] This step is based on linear frequency modulation signals. A pilot section that can be accurately monitored is designed by using a Hamming window to eliminate spectral leakage, so as to facilitate the accurate segmentation of the sensing part in the future.
[0010] Step 1.2: Design of the perception part.
[0011] This step takes into account both avoiding spectrum leakage and preserving a sufficient amount of effective spectrum. The sensing part is designed by applying Hamming windows to the front and rear regions of the linear frequency modulated signal in order to fully preserve the propagation characteristics of the acoustic signal within the mobile phone's internal structure.
[0012] Step 1.3: Audio signal parameter settings.
[0013] This step takes into account the limitations of environmental noise distribution and frequency selection in high-frequency regions to set the frequency range of the sensing signal; it takes into account the signal-to-noise ratio and system response delay to set the signal duration; and in order to avoid the sensing part from interfering with the detection of the pilot part, it selects to use down-modulation and up-modulation linear frequency modulation signals to modulate the pilot part and the sensing part.
[0014] Step 2: Signal segmentation.
[0015] Step 2.1: Environmental noise elimination
[0016] This step utilizes bandpass filtering, combined with the start and end frequencies of the sensing signal set in step 1, to eliminate environmental noise from the collected active audio sensing samples.
[0017] Step 2.2: Signal synchronization and segmentation.
[0018] This step first combines cross-correlation and peak detection methods to accurately determine the starting position of the leading signal in the audio sample. Then, using this position as the starting point, the audio segments of the perceived portion are segmented according to the temporal structure of the perceived signal in step 1.
[0019] Step 3: Reference signal generation.
[0020] This step takes the active audio perception samples collected in outdoor non-handheld scenarios, processes them in step 2, and generates a reference signal by frequency domain averaging, which is used for subsequent structured sound signal extraction.
[0021] Step 4: Feature extraction.
[0022] Step 4.1: Elimination of direct air path interference.
[0023] This section first performs frequency domain averaging on the signal segments corresponding to the perceived portion obtained after processing the samples collected in the user-held scenario in step 2. Then, the reference signal generated in step 3 is eliminated in the frequency domain to remove direct air path interference and obtain the structure-borne acoustic signal.
[0024] Step 4.2: Time-frequency feature extraction.
[0025] This step uses short-time Fourier transform to process the structured acoustic conduction signal obtained in step 4.1 to obtain time-frequency information. Then, by cropping out the portion outside the sensing frequency band, time-frequency features are generated.
[0026] Step 5: Model training and certification.
[0027] Step 5.1: Model training.
[0028] In the registration phase, this step trains the model on the time-frequency features extracted from user samples during registration to enhance feature differences between users. Then, the reconstructed features corresponding to each sample output by the CNN model are averaged and stored in the model library as user models.
[0029] Step 5.2: Identity Authentication.
[0030] In the authentication phase, this step extracts the time-frequency features from the samples input by the user to be authenticated and reconstructs these features based on the CNN model used in the registration phase. Combining similarity and threshold methods, it identifies whether the user to be authenticated is a registered or unregistered user.
[0031] Beneficial effects
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] 1. This invention utilizes only the microphones and speakers commonly integrated in commercial smartphones to achieve a fast, hardware-free, and widely applicable continuous authentication method. This invention leverages the structural acoustic signal transmission correlation effect of the phone, extracting the characteristics of this sound signal to characterize the user's hand-holding posture-related biometrics, effectively resisting imitation attacks.
[0034] 2. This invention utilizes active audio sensing samples in outdoor non-handheld scenarios to generate reference signals, thereby eliminating signal interference from direct air propagation paths and improving the reliability of identity authentication.
[0035] 3. The continuous identity authentication method involved in this invention is effective and robust, achieving an average authentication accuracy of approximately 97% in an identity verification experiment involving 20 volunteers. Attached Figure Description
[0036] Figure 1 This is a diagram illustrating the architecture of the continuous identity authentication method according to an embodiment of the present invention.
[0037] Figure 2 This is a diagram of the sensing signal structure according to an embodiment of the present invention.
[0038] Figure 3 CNN network structure
[0039] Figure 4 For the overall performance of the embodiments of the present invention
[0040] Figure 5 To improve the anti-mimicry attack performance of embodiments of the present invention Detailed Implementation
[0041] The present invention will now be described in further detail with reference to the embodiments and accompanying drawings.
[0042] A persistent authentication method for smartphones based on structured acoustic signals involves two stages: registration and authentication (e.g., Figure 1(As shown). Before the registration and authentication phases, this invention pre-designs an acoustic sensing signal based on linear frequency modulation (LFM) signals (consisting of a leader, a sensing part, and an idle part) for active audio sensing. During the registration phase, users need to collect samples in both outdoor non-handheld and handheld scenarios. Samples collected in outdoor non-handheld scenarios undergo environmental noise cancellation, signal synchronization, and segmentation; the segmented sensing part serves as a reference signal for eliminating direct propagation path signals. Samples collected in handheld scenarios undergo signal segmentation, direct propagation path signal elimination, and time-frequency feature extraction before being used for model training; the corresponding model is stored in the model library. During the authentication phase, users hold the device using the same hand posture as in the registration phase and collect samples. The collected samples undergo environmental noise cancellation and signal segmentation, and structural conduction audio features are extracted. These are then combined with the model library generated in the registration phase for user authentication. The process includes the following steps:
[0043] Step 1: Sensing signal generation.
[0044] This invention selects the bottom speaker and top microphone of a smartphone to perform active audio sensing, ensuring that the sensing signal propagation path covers the entire body of the phone. The acoustic sensing signal used in this invention (such as...) Figure 2 (As shown) It consists of a pilot section, a sensing section, and an idle section. The idle section is designed to prevent interference between preceding and following signals.
[0045] The design of the pilot section and the sensing section is as follows:
[0046] Step 1.1 Design of the pilot section.
[0047] The lead-in portion is primarily used to accurately determine its starting position within the audio stream recorded by the microphone, enabling accurate segmentation of the sensing portion. Since linear frequency modulated (LFM) signals possess high autocorrelation, this invention modulates the lead-in portion based on an LFM signal. The definition of the LFM real-valued signal s(t) is as follows:
[0048]
[0049] Where A is the signal amplitude, T is the signal duration, and the frequency slope is... B = f1 - f0 is the frequency bandwidth, and the center frequency is... Where f0 and f1 are the start and end frequencies of the linear frequency modulated signal, respectively. The phase at t=0.
[0050] To reduce the impact of spectral leakage on accurate detection of the leader, this invention applies a Hamming window to the entire linear frequency modulated signal s(t), i.e.:
[0051]
[0052] in,
[0053]
[0054] Step 1.2 Design of the sensing component
[0055] The sensing component is used to extract structurally conducted sound signal features. Structural acoustic conduction exhibits a frequency selectivity effect; that is, when sound signals propagate inside a mobile phone, different frequencies of sound waves exhibit varying conduction capabilities due to the structural physical characteristics. This results in some frequencies of sound propagating more easily through the structure, while others are suppressed or attenuated. Therefore, the sensing component needs a wide bandwidth to effectively extract structurally conducted sound signal features. Thus, this invention also uses a linear frequency modulated signal s(t) to modulate the sensing component. On one hand, the original linear frequency modulated signal s(t) can be considered as having a square window applied, where the entire frequency band amplitude is unmodulated. However, signal truncation leads to frequency leakage. On the other hand, applying a Hamming window to the entire time domain of s(t) can effectively avoid frequency leakage, but the Hamming window weights both ends of the signal, causing a decrease in the total signal energy and reducing the effective bandwidth. Therefore, this invention chooses a compromise approach to avoid frequency leakage while ensuring a sufficient effective bandwidth, i.e., modulating the front and back ends of the s(t) signal. Apply rising and falling edge Hamming windows to the intervals respectively:
[0056]
[0057] in,
[0058]
[0059] Step 1.3 Audio signal parameter settings.
[0060] Commercial smartphones generally support an acoustic sampling rate of up to 48kHz. To prevent aliasing interference during recording, a 24kHz low-pass filter is applied to the microphone. Furthermore, frequencies below 8kHz are easily affected by ambient noise. Considering that speakers and microphones generally have low response times above 23kHz, this invention sets the frequency band for the pilot and sensing components to 8kHz-23kHz.
[0061] The longer the duration of an audio signal, the higher its signal-to-noise ratio, but this also increases perceptual latency. If the duration of the perceived signal is short, the frequency resolution will be reduced.
[0062] In the active audio perception signal, the pilot signal is played only once. To ensure its signal-to-noise ratio and improve the detection accuracy of the pilot signal, the present invention empirically sets the duration of the pilot signal to 0.2 seconds. To ensure the frequency resolution and perception delay of the perception part, the present invention sets the duration of the perception part to 0.1 seconds. The idle interval is set to 0.15 seconds. To improve the stability of the perception part, three perception parts are set in the active audio perception signal. In addition, the present invention uses the orthogonality of the chirp signal to avoid the interference of the perception part on the detection of the pilot part. That is, as Figure 2 shown, the present invention modulates the pilot part and the perception part with the down-chirp and up-chirp chirp signals respectively.
[0063] Step 2: Signal segmentation.
[0064] This part eliminates the environmental noise from the audio sample r(i), i∈[1,N] recorded by the microphone and performs segmentation to obtain the audio segment corresponding to the perception part in the preset perception signal. The main steps are as follows:
[0065] Step 2.1: Apply the Butterworth band-pass filtering method BandFilter(·) to filter the audio sample r(i) according to the preset start frequency f0 and end frequency f1 of the perception signal in Step 1 to obtain the audio signal after eliminating the environmental noise That is:
[0066]
[0067] Step 2.2: Apply the cross-correlation operation xcorr(·,·) to determine the starting position i of the pilot signal PilotSignal in the sample signal p . For example, for the cross-correlation operation R = xcorr(s1, s2) of the s1 signal (length L) and the s2 signal (length l, and l < L), its i-th element R(i), i∈[1, L + l - 1] represents the similarity between the s1 signal and the s2 signal delayed by i sample lengths. Therefore, the sample corresponding to the peak in the output result R of the cross-correlation operation is the starting position of the s2 signal in the s1 signal. Therefore, by performing peak detection PeakFind(·) on R, the starting sample position i of the PilotSignal in the signal can be obtained P . That is:
[0068]
[0069] Then, starting from the position i p , the starting sample positions of the subsequent three perception parts are determined in turn according to the durations of the pilot part, the idle part and the perception part in Step 1 x = 1, 2, 3. That is:
[0070]
[0071] Among them, L P L S With L I These are the sample lengths corresponding to the leading part, the sensing part, and the idle part in the preset sensing signal, respectively.
[0072] In addition, in order to eliminate ambient noise from the audio signal The present invention completely extracts the structural transmission signal, and the actual extraction start position is from the starting position of the sensing part. Shift the sample position to the left by offset (offset = 200). That is, the present invention... In the signal, with Starting from the sample index position, the truncation length is (L) S The audio segment with +offset is used as the audio segment of the perception part. Right now:
[0073]
[0074] Step 3: Reference Signal Generation
[0075] During the registration phase, samples collected from users in outdoor, non-handheld scenarios are used to generate reference signals for subsequent elimination of direct propagation path signals. In outdoor, non-handheld scenarios, the multipath effect is less pronounced than in indoor scenarios, so the influence of multipath signals in the collected samples can be ignored. Furthermore, although samples collected in non-handheld scenarios still contain sound signals transmitted through the phone's internal structure, this portion of the signal is not affected by the user's holding posture and can therefore be used as a reference signal for eliminating airborne direct propagation signals. The specific steps are as follows: Assume the perceived audio segment obtained from the samples collected in outdoor, non-handheld scenarios after processing in step 2 is Ref x x = 1, 2, 3. Then, the reference signal Ref passes through Ref... x Obtained by averaging in the frequency domain. That is:
[0076] Ref = IFFT(AVE(FFT(Ref) x x = 1, 2, 3
[0077] Here, AVE is an operation that averages multiple one-dimensional horizontal quantities column-wise. FFT and IFFT are the Fast Fourier Transform and Inverse Fast Fourier Transform, respectively.
[0078] Step 4: Feature extraction.
[0079] Step 4.1 Elimination of Direct Air Path Interference. Compared to signals propagating directly through the air, structured sound signals propagating inside the device, although with a shorter propagation distance (approximately the length of the phone's body), experience greater loss due to solid-media propagation. Therefore, samples collected in user-held scenarios show stronger energy from direct air path propagation, interfering with the propagation characteristics of structured sound signals. Thus, it is necessary to eliminate the influence of direct air path propagation.
[0080] Assuming the sample collected in the user holding scenario is the perceived audio segment obtained in step two, let's call it Sam. x x = 1, 2, 3. Similar to the generation of the reference signal, a more stable sample Sam is obtained through formula (3): Sam = IFFT(AVE(FFT(Sam) x ))). Here, AVE is an operation that averages multiple one-dimensional horizontal quantities column-wise. FFT and IFFT are the Fast Fourier Transform and Inverse Fast Fourier Transform, respectively. This invention utilizes frequency domain subtraction to eliminate air propagation path interference, that is:
[0081]
[0082] Step 4.2: Time-frequency feature extraction.
[0083] This invention uses STFT (Short Time Fourier Transform) to obtain The time-spectrum information of the signal is obtained. Then, it is cropped based on the audio frequency band range [f0, f1] in step 1 to reduce the computational load.
[0084] Step 5: Model training and certification.
[0085] Step 5.1 Model Training. This invention uses a CNN (Convolutional Neural Network) model to train the features extracted from registered samples to enhance the feature differences between users. The specific CNN structure used in this invention is as follows: Figure 3 As shown. The input layer of this CNN is the time-frequency map feature generated in step 4.2. The CNN structure includes three convolutional blocks, one fully connected (FC) layer, one flattened layer, and one FC fully connected layer. Each convolutional block includes two convolutional layers (Conv2D+BN+ReLU), one MaxPooling layer, and one Dropout layer. The Conv2D layer and MaxPooling layer use 3×3 and 2×2 kernels, respectively. At this point, the average value of the output features of the FC layer in the CNN corresponding to each user sample is used as the model for each user and stored in the model library.
[0086] Step 5.2 Identity Authentication: In the authentication phase, the features corresponding to the input sample are used to reconstruct the features through the CNN model from the registration phase. Then, the similarity between the stored models is calculated using Euclidean distance. The user to be authenticated is identified as the registered user with the highest similarity. If the similarity is lower than a preset threshold (0.85), the user is identified as an unregistered user.
[0087] Example
[0088] The authentication system involved in this invention is implemented based on an Android mobile device (Redmi K40) and Matlab. The Android mobile device is responsible for user sample collection. Matlab runs on a PC and is responsible for subsequent signal processing, feature extraction, model training, and authentication. Sample data was collected in three scenarios: outdoors, in a laboratory, and in a conference room. Samples collected in the outdoor scenario were used to generate reference signals, while samples collected in the laboratory (relatively quiet) and conference room (noisy with music playing) were used for user authentication system performance evaluation and attack resistance verification. Twenty volunteers (8 females and 12 males, aged 21 to 48) were recruited for each of the laboratory and conference room scenarios to participate in sample data collection. Fifteen volunteers played the role of legitimate users, and the remaining five played the role of attackers. Specifically, during the registration phase, each volunteer playing the role of a legitimate user collected 20 samples each in the laboratory and conference room while maintaining a comfortable grip.
[0089] Authentication accuracy
[0090] First, the overall authentication accuracy of the authentication method involved in this invention is evaluated. All samples from the five attackers were not used in training but directly participated in testing. Therefore, it can be considered that these five attackers carried out a random attack. Figure 3 This is a confusion matrix showing the authentication results of 15 legitimate users and 5 attackers (labeled F) according to the present invention. The figure indicates that the successful authentication accuracy for legitimate users is 97.3%, and it can identify attackers carrying out random attacks with an accuracy of 99.6%. This demonstrates that the authentication method of the present invention can effectively and accurately identify legitimate users and attackers carrying out random attacks.
[0091] Effectiveness of mimic attack defense
[0092] To verify whether this invention can effectively resist imitation attacks, this section involves five attackers each selecting a legitimate registered user's grip posture for thorough observation. Then, they attempt to log in to the system by imitating the legitimate user's grip posture. During the model training phase, 10 and 20 samples from legitimate users were used respectively to verify the impact of changes in the training samples on the effectiveness of resisting imitation attacks. Figure 4The false acceptance rates (FACKs) for attacking users attempting authentication are shown in laboratory and conference room scenarios. With 20 training samples, the FACKs in both laboratory and conference room scenarios are below 0.15%. This demonstrates that the authentication method of this invention effectively resists imitation attacks. Furthermore, with 10 training samples, the FACK rate improves, but remains below 0.25%. This indicates that the authentication method of this invention is more effective at resisting imitation attacks when users provide more registration samples.
Claims
1. A method for continuous identity authentication of smartphones based on structure-conducted acoustic signals, characterized in that: Step 1: Generating sensing signals, as detailed below; Step 1.1: Pilot section design; This step is based on a linear frequency modulated signal, and the pilot section is designed using a Hamming window to eliminate spectral leakage; Step 1.2: Design of the sensing component; The sensing section was designed by applying Hamming windows to the front and rear regions of the linear frequency modulated signal; Step 1.3: Audio signal parameter settings; This step takes into account the limitations of environmental noise distribution and frequency selection in high-frequency regions to set the frequency domain range of the sensing signal. Step 2: Signal segmentation; Step 2.1: Environmental noise elimination This step utilizes bandpass filtering, combined with the start and end frequencies of the sensing signal set in step 1, to eliminate environmental noise from the collected active audio sensing samples. Step 2.2: Signal synchronization and segmentation; This step first combines cross-correlation and peak detection methods to accurately determine the starting position of the leading signal in the audio sample; then, using this position as the starting point, the audio segments of the perception part are segmented. Step 3: Reference signal generation; This step generates a reference signal by averaging the active audio perception samples collected in outdoor non-handheld scenarios in the frequency domain. Step 4: Feature extraction; Step 4.1: Elimination of direct air path interference; This section first performs frequency domain averaging on the signal segments corresponding to the sensing part obtained after processing the samples collected in the user holding scenario in step 2. Then, the reference signal generated in step 3 is eliminated in the frequency domain to eliminate direct air path interference and obtain the structural acoustic conduction signal; Step 4.2: Time-frequency feature extraction; This step uses short-time Fourier transform to process the structured acoustic transmission signal obtained in step 4.1 to obtain time-frequency information; then, by cropping out the part outside the sensing frequency band, time-frequency features are generated. Step 5: Model training and certification; Step 5.1: Model training; In the registration phase, this step trains the time-frequency features extracted from user samples during the registration phase based on a convolutional neural network (CNN) model to enhance the feature differences between users. Then, the reconstructed features corresponding to each sample output by the CNN model are averaged and stored in the model library as the user model. Step 5.2: Identity Authentication; In the authentication phase, this step extracts the time-frequency features of the samples input by the user to be authenticated and reconstructs the features based on the CNN model from the registration phase. Combining similarity and threshold methods, it identifies whether the user to be authenticated is a registered user or an unregistered user.
2. The method according to claim 1, characterized in that... Includes the following steps: Step 1: Generating sensing signals; Step 1.1 Design of the pilot section; The pilot portion is modulated based on a linear frequency modulated signal; Linear frequency modulated real signal The definition is as follows: in, The signal amplitude, The duration of the signal and the frequency slope , For frequency bandwidth, center frequency ;in and These are the start and end frequencies of the linear frequency modulated signal, respectively. for Phase of time; To reduce the impact of spectral leakage on accurate detection of the leader, the linear frequency modulated signal is subjected to... The entire structure includes Hanming windows, namely: in, Step 1.2 Design of the Sensing Component Using linear frequency modulation signal Modulate the sensing part; that is, modulate the sensing part. Signal front and back ends Apply rising and falling edge Hamming windows to the intervals respectively: in, Step 1.3 Audio signal parameter settings; The frequency bands for the pilot section and the sensing section are set to 8kHz-23kHz; Set the pilot signal duration to 0.2 seconds; set the sensing duration to 0.1 seconds; and set the idle interval to 0.15 seconds. Step 2: Signal segmentation; This section contains audio samples recorded by the microphone. The process involves eliminating environmental noise and segmenting the signal to obtain the audio segment corresponding to the perceived portion of the preset sensing signal. This mainly includes the following steps: Step 2.1: Apply the Butterworth bandpass filtering method Based on the preset sensing signal start frequency in step 1 With termination frequency For audio samples Filtering is performed to obtain the audio signal after eliminating ambient noise. ;Right now: Step 2.2: Apply cross-correlation operations Determine the sample signal Middle pilot signal starting position ;for Signals and Cross-correlation operation of signals , its first element Indicates signal Signal and Delay Sample length The similarity between signals; therefore, the cross-correlation operation outputs the result. The sample corresponding to the mid-peak value is The signal is The starting position in the signal; therefore, by... Perform peak detection Get exist Starting sample position in the signal ;Right now: Then with Starting from the current location, the initial sample locations for the subsequent three sensing parts are determined sequentially according to the durations of the pilot part, idle part, and sensing part in step 1. ;Right now: in, , and These are the sample lengths corresponding to the leading part, sensing part, and idle part of the preset sensing signal, respectively. To obtain the audio signal after eliminating ambient noise The structural transmission signal is completely intercepted, and the actual interception start position is set from the starting position of the sensing part. Move to the left One sample ( =200); that is, in In the signal, with ( The sample index position is the starting point, and the truncation length is... Audio segments as part of the perception audio segment ;Right now: Step 3: Reference Signal Generation The perceptual audio segments obtained from samples collected in outdoor non-handheld scenarios after processing in step 2 are as follows: Therefore, the reference signal Through the Obtained by frequency domain averaging; that is: in, The operation of averaging multiple one-dimensional horizontal quantities by column; and These are the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively. Step 4: Feature extraction; Step 4.1 Elimination of direct air path interference; The samples collected in the user-held scenario, after being processed through step two to obtain the perceived audio segments, are... To obtain more stable samples = ;in, The operation of averaging multiple one-dimensional horizontal quantities by column; and These are the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively; interference from the air propagation path is eliminated by subtracting data in the frequency domain, i.e.: Step 4.2: Time-frequency feature extraction; Obtain using Short Time Fourier Transform (STFT) The time-spectrum information of the signal; then, based on the audio frequency band range in step 1. Cut; Step 5: Model training and certification; Step 5.1 Model Training: The features extracted from the registered samples are trained using a Convolutional Neural Network (CNN) model to enhance the feature differences between users; the average value of the output features of the FC layer in the CNN corresponding to each user sample is used as the model for each user and stored in the model library. Step 5.2 Identity Authentication: In the authentication phase, the features corresponding to the input sample are used to output reconstructed features through the CNN model in the registration phase; then, the similarity of the stored models is calculated using Euclidean distance; the user to be authenticated is identified as the registered user with the highest similarity; if the similarity is lower than the preset threshold of 0.85, it is identified as an unregistered user.
Citation Information
Patent Citations
Keystroke and sound signal fused user non-inductive credible identity authentication method, system and terminal
CN116910732A
Implicit answering authentication method and system based on external ear acoustic perception
CN117156439A