PIN code identity authentication method and system based on acoustic structure response

By collecting and analyzing the sound of users clicking their PIN codes, and using acoustic structural response and pre-trained audio neural networks for feature extraction and classification, the problem of PIN codes being vulnerable to shoulder spying attacks is solved. This achieves high-precision identity authentication and security decoupling, improving the security of mobile devices and the user experience.

CN121865273APending Publication Date: 2026-04-14HUNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing PIN code authentication mechanisms are vulnerable to shoulder spying attacks. The input process is highly coupled with the verification content, causing the identity authentication system to fail. Traditional protection methods are difficult to promote on a large scale in consumer devices.

Method used

By collecting user PIN code click sound samples, identity authentication is performed using acoustic structure response. Pre-trained audio neural networks (PANNs) with sound field LESR enhancement are used for feature extraction and classification, thereby decoupling authentication content from authentication criteria.

Benefits of technology

It effectively protects the PIN code input process, prevents shoulder spying attacks, achieves high-precision user identification, allows arbitrary number input, and the system cannot be imitated or replayed to impersonate the user, thus improving security and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865273A_ABST
    Figure CN121865273A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of identity authentication of mobile terminals, and provides a PIN code identity authentication method and system based on acoustic structure response, and the method comprises the steps: collecting a PIN code click sound sample when a user carries out identity authentication; fine-grained denoising is carried out on the PIN code click sound sample to obtain anti-noise biological features, and the fine-grained denoising comprises key frequency band selection and normalized segmentation; performing feature extraction on the anti-noise biological features by using a sound field LESR enhanced pre-trained audio neural network PANNs to obtain an acoustic feature extraction result; and the acoustic feature extraction result is classified by adopting a single-sample classification and multi-sample authentication mode to obtain a PIN code identity authentication result, so that the security of unlocking the mobile terminal by using the PIN code by a user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mobile terminal identity authentication technology, and in particular to a PIN code identity authentication method and system based on acoustic structural response. Background Technology

[0002] As the core carrier of personal information and sensitive data, smartphones offer great convenience but also bring serious privacy and security risks. Identity authentication, a key step in mitigating these risks, still primarily relies on PIN codes—their advantages lie in their simplicity, ease of memorization, and recovery capabilities, making them widely used for device unlocking and application access. Although biometric technologies such as Face ID and fingerprint recognition have become increasingly common, many users still depend on PIN codes; even in systems that support biometrics, PIN codes are often used as a backup mechanism in case of authentication failure, highlighting their enduring value as a security safety net.

[0003] However, PIN code authentication mechanisms also have significant drawbacks: the input process is inherently observable, making them highly vulnerable to "shoulder spying attacks"—attackers can directly obtain the user's input sequence of numbers through visual observation or video surveillance, thereby stealing the PIN code. How to effectively protect the PIN code input process from shoulder spying attacks remains a major challenge that urgently needs to be addressed in the field of mobile identity authentication. Commonly used defensive measures, such as tilting the screen to obstruct the view, offer only limited protection and are easily bypassed by tools such as mirrors and hidden cameras.

[0004] Existing technical solutions can be mainly divided into three categories:

[0005] 1) Visual obfuscation technology: Through dynamic masking, character deformation and other methods, the input content is made visible only to the user and not to onlookers;

[0006] 2) Environmental sensing technology: using sensors (such as infrared, ultrasonic, or cameras) to detect whether there are potential spyers in the surroundings;

[0007] 3) Physical shielding technology: Using screen dimming, directional filters, or hardware-level privacy films to limit the viewing angle and reduce information leakage. Although the above solutions have certain protection potential, they still face significant limitations in practical applications: the first type of solution can usually only delay attacks and cannot completely block them; the second and third types of solutions often rely on additional hardware support, or lead to reduced screen visibility and impaired user experience, making it difficult to promote them on a large scale in consumer devices.

[0008] Therefore, the security bottleneck of traditional PIN code authentication mechanisms lies in the high coupling between user input and verification content; each keystroke is directly mapped to a digital symbol that can be observed externally. Once the PIN code is stolen, the system will be unable to distinguish between legitimate users and imposters, thus rendering the identity authentication system ineffective. Summary of the Invention

[0009] Aimed at at least in solving one of the technical problems existing in the prior art, the present invention provides a PIN code authentication method and system based on acoustic structural response.

[0010] One aspect of the present invention provides a PIN code authentication method based on acoustic structural response, comprising:

[0011] Collect sound samples of PIN code clicks during user authentication;

[0012] Fine-grained denoising is performed on the PIN code click sound sample to obtain noise-resistant biometric features, wherein fine-grained denoising includes key frequency band selection and normalized segmentation;

[0013] Pre-trained audio neural networks (PANNs) with enhanced sound field LESR were used to extract the noise-resistant biofeedback features, resulting in acoustic feature extraction.

[0014] The acoustic feature extraction results are classified using single-sample classification and multi-sample authentication methods to obtain PIN code authentication results.

[0015] According to the PIN code authentication method based on acoustic structure response, the method collects PIN code click sound samples when a user authenticates their identity, including:

[0016] Obtain PIN code click sound samples captured by the top and bottom microphones of the mobile device:

[0017]

[0018] in The path-dependent channel frequency response; Represents frequency Click spectrum at the location; Represents environmental noise signals; for:

[0019]

[0020] in, Indicates the distance from the clicked location to the microphone. The effective transmission distance; Indicates geometric diffusion; This indicates the structural damping and scattering within the fuselage; This represents the finger-surface contact impedance and receiver chain gain.

[0021] 3. The PIN code authentication method based on acoustic structural response according to claim 2, characterized in that the method further includes:

[0022] Click the PIN code to analyze the frequency band of the sound sample. The short-time Fourier transform is used to uniformly divide the area into... A continuous sub-band ;

[0023] Based on the sub-band responses from the top and bottom microphones, determine the PIN code click sound sample in the sub-band. and the LESR features of frames for:

[0024]

[0025] in, For the child response of the top microphone; For the sub-band response of the bottom microphone; and To fix the click position, and This is a band-limited parameter.

[0026] According to the PIN code authentication method based on acoustic structural response, the key frequency band selection includes:

[0027] Perform L-level stationary wavelet transform on the PIN code click sound samples from the top and bottom microphones to obtain frequency band components at different stationary wavelet transform levels, and then divide the subbands... By mapping the center frequency to the corresponding stationary wavelet transform level, a complete LESR feature matrix with sub-bands and frames is obtained, containing a set of mapping relationships. for:

[0028]

[0029] in, This serves as an identifier for the stationary wavelet transform level. For children The center frequency; The passband corresponds to the stationary wavelet transform level; the complete LESR characteristic matrix is:

[0030]

[0031] Stability scoring of stationary wavelet transform levels:

[0032]

[0033] in, Indicates the first The scoring results of the level of stationary wavelet transform. For subband stability;

[0034] The stationary wavelet transform levels are sorted according to the stability score. A preset number of stationary wavelet transform levels are selected from the sorting results and inverse stationary wavelet transform is performed to obtain the LESR enhanced signal.

[0035] According to the PIN code authentication method based on acoustic structural response, the normalized segmentation includes:

[0036] By acquiring the sound field asymmetry difference between the tapping events during user authentication and the ambient noise signal, the overall... parameter for:

[0037]

[0038] in, for The full-band energy of the top microphone channel at any given moment; for Full-band energy of the bottom microphone channel at all times; Used to characterize the degree of signal asymmetry between the channels of the top microphone and the channels of the bottom microphone;

[0039] Calculate the whole parameter The sample mean is used to select candidate tapping event frames from the sample mean by using a preset judgment threshold.

[0040] According to the PIN code authentication method based on acoustic structural response, pre-trained audio neural networks (PANNs) with enhanced acoustic field LESR are used to extract features from the noise-resistant biometrics, resulting in acoustic feature extraction results, including:

[0041] The PIN code click sound samples were extracted using pre-trained audio neural networks (PANNs) to obtain high-frequency time-frequency features;

[0042] High-frequency time-frequency features and noise-resistant biometric features are processed using a lightweight fusion network to obtain a fused embedding;

[0043] The fusion embedding is decoupled using contrastive learning to obtain acoustic feature extraction results, where the total loss function of contrastive learning is... for

[0044]

[0045] The interval-based triplet loss function, for The weight parameters, The calculation method is as follows:

[0046]

[0047] in, For integration and embedding, This represents anchor samples, positive samples, and negative samples. Positive samples... With anchor point sample The same users, negative samples With anchor point sample Different users;

[0048] The binary cross-entropy loss function is... for The weight parameters, The calculation method is as follows

[0049]

[0050] in, It is the logit value predicted by the model. It's a label indicating whether the current sample is valid. These are the weight parameters.

[0051] According to the PIN code authentication method based on acoustic structural response, the acoustic feature extraction results are classified using single-sample classification and multi-sample authentication methods to obtain PIN code authentication results, including:

[0052] The acoustic feature extraction results are converted into confidence scores using the sigmoid function. Among them, the confidence score Used to characterize the probability that the PIN code click sound sample comes from the target user;

[0053] The single-sample classification is based on confidence scores. Use preset thresholds to determine whether a single tap in the PIN code click sound sample comes from a legitimate user;

[0054] The multi-sample authentication is based on confidence scores. The session-level acceptance probability is used to determine whether multiple PIN code click sounds originate from a legitimate user. This session-level acceptance probability includes the legitimate user session acceptance probability and the attacker session acceptance probability. The legitimate user session acceptance probability is:

[0055]

[0056] The attacker's session acceptance probability is:

[0057]

[0058] in, The number of keystrokes required for authentication. The number of times the PIN code was tapped to produce the sound sample. For the single-sample authentication success rate of legitimate users, The attacker's single-sample false positive rate. This is used as a sample serial number identifier.

[0059] According to the PIN code authentication method based on acoustic structure response, the preset threshold is determined based on the F1 score and error rate of the evaluation index.

[0060] Another aspect of the present invention provides a PIN code authentication system based on acoustic structural response, comprising:

[0061] The first module is used to collect sound samples of PIN code clicks when users authenticate their identity.

[0062] The second module is used to perform fine-grained denoising on the PIN code click sound sample to obtain noise-resistant biometric features, wherein fine-grained denoising includes key frequency band selection and normalized segmentation.

[0063] The third module is used to extract features from the noise-resistant biofeedback using pre-trained audio neural networks (PANNs) enhanced with sound field LESR, and obtain acoustic feature extraction results.

[0064] The fourth module is used to classify the acoustic feature extraction results using single-sample classification and multi-sample authentication methods to obtain PIN code authentication results.

[0065] The beneficial effects of this invention are as follows: It utilizes the natural sound waves generated by a user's touchscreen input during PIN code entry to construct an identity authentication mechanism based on acoustic biometrics. The core principle is that when a user's finger touches the screen, the generated acoustic signal propagates through the device's internal structure and is simultaneously collected by multiple embedded microphones distributed in different locations on the device (such as the top earpiece microphone and the bottom main microphone). Because the sound waves are affected by materials, structure, and contact methods during propagation, path-dependent frequency attenuation and resonance characteristics occur, thus naturally amplifying individual microbial differences among users.

[0066] By modeling and analyzing the aforementioned multi-channel acoustic signals, the system can achieve high-precision user identification in complex environments (such as background noise and different grip postures). More importantly, this mechanism allows users to input arbitrary numbers (including random or "bait" PIN codes), and the system verifies identity solely based on acoustic behavioral characteristics, thus achieving complete decoupling between authentication content and authentication criteria. In this way, even if an attacker completely observes the input number sequence, they cannot impersonate the user through imitation or replay, fundamentally mitigating the security risks caused by PIN code leakage. Attached Figure Description

[0067] Figure 1 This is a schematic diagram of the PIN code authentication process based on acoustic structural response according to an embodiment of the present invention.

[0068] Figure 2 This is a basic architecture diagram of the identity authentication framework in an embodiment of the present invention.

[0069] Figure 3 This is a diagram of the PANNS-based feature extraction model according to an embodiment of the present invention.

[0070] Figure 4 This represents the probability of accepting a legitimate user session under different N and M values ​​in this embodiment of the invention.

[0071] Figure 5 These are the recall rate and 1-false positive rate of different users' sessions in this embodiment of the invention.

[0072] Figure 6 These are the session acceptance recall rate and 1-false positive rate under different environments in this embodiment of the invention.

[0073] Figure 7 These are the recall rate and 1-false positive rate of session acceptance at different times in this embodiment of the invention.

[0074] Figure 8 This is a schematic diagram of a PIN code authentication system based on acoustic structural response, according to an embodiment of the present invention. Detailed Implementation

[0075] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings. Throughout the description, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions. In the following description, suffixes such as "module," "part," or "unit" used to denote elements are used only for the purpose of illustrative purposes and have no specific meaning in themselves. Therefore, "module," "part," or "unit" can be used interchangeably. Terms such as "first," "second," etc., are used only to distinguish technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the sequential relationship of the indicated technical features. In the following description, the consecutive reference numerals for method steps are for ease of review and understanding. Adjusting the implementation order of steps, in conjunction with the overall technical solution of the present invention and the logical relationship between the various steps, will not affect the technical effect achieved by the technical solution of the present invention. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0076] refer to Figure 1 , Figure 1 This is a schematic diagram of the PIN code authentication process based on acoustic structural response according to an embodiment of the present invention, which includes, but is not limited to, steps S100~S400:

[0077] S100 collects sound samples of PIN code clicks during user authentication.

[0078] Understandably, when a user taps the PIN code, the resonance between the user's finger and the phone screen generates sound. This sound signal reflects and propagates multiple times within the phone's internal structure, eventually reaching and being captured by the phone's top and bottom microphones. This sound is then transmitted through the phone screen to the top and bottom microphones.

[0079] In some embodiments, PIN code click sound samples collected by the top and bottom microphones of the mobile terminal are obtained:

[0080]

[0081] in The path-dependent channel frequency response; Represents frequency Click spectrum at the location; Represents environmental noise signals; for:

[0082]

[0083] in, Indicates the distance from the clicked location to the microphone. The effective transmission distance; Indicates geometric diffusion; This indicates the structural damping and scattering within the fuselage; This represents the finger-surface contact impedance and receiver chain gain.

[0084] It should be noted that, and frequency in Originating from different signal sources, this decomposition method decomposes channel-related effects. (Determined by attenuation and coupling) and effects related to the signal source. (The click behavior is encoded) Separate it.

[0085] In some embodiments, in order to formally capture the frequency domain amplitude differences in the device's sound field caused by structure propagation, the present invention analyzes the frequency band of the PIN code click sound sample. The short-time Fourier transform is used to uniformly divide the area into... A continuous sub-band For a clean signal, the noise term in the formula... It can be ignored. Indicates length is The frame index within the click time interval of the frame. Then, in this embodiment of the invention, the subband... and frame place Defined as:

[0086]

[0087] in and These represent the sub-band responses of the top and bottom microphones, respectively. This is because both channels share the same source. Therefore, this term cancels out in the ratio, making Depends only on the relative transfer function and Only a tiny amount of noise remains.

[0088] Conversely, for environmental noise signals (which does not exist) ), This only reflects the noise components captured by the two microphones. In general, according to formula (2), the embodiments of the present invention typically derive...

[0089]

[0090] For ambient noise, the propagation path to both microphones is actually the same, therefore Naturally tending towards zero. Conversely, for a fixed click location, and band-limited parameters Remain unchanged, where This indicates whether it's a top microphone or a bottom microphone, which ensures... It exhibits high stability during repeated PIN code inputs. Furthermore, it is inherently robust to variations in click pressure and user behavior, providing a basis for detecting replay attacks that are noise-resistant and aware of sound sources by leveraging sound field effects.

[0091] S200 performs fine-grained denoising on PIN code click sound samples to obtain noise-resistant biometric features. The fine-grained denoising includes key frequency band selection and normalized segmentation.

[0092] Understandably, when a user enters a PIN code, the captured audio signal is a mixture of genuine click responses and ambient noise, which can interfere with user authentication based on biometrics embedded in the signal. Therefore, it is necessary to perform fine-grained denoising on the signal while preserving its inherent biometric characteristics.

[0093] In some embodiments, to achieve effective signal enhancement, the present invention provides a key frequency band selection method based on multi-resolution analysis as follows:

[0094] Perform L-level stationary wavelet transform on the PIN code click sound samples from the top and bottom microphones to obtain frequency band components at different stationary wavelet transform levels, and then divide the subbands... By mapping the center frequency to the corresponding stationary wavelet transform level, a complete LESR feature matrix with sub-bands and frames is obtained, containing a set of mapping relationships. for:

[0095]

[0096] in, This serves as an identifier for the stationary wavelet transform level. For children The center frequency; The passband corresponds to the stationary wavelet transform level; the complete LESR characteristic matrix is:

[0097]

[0098] Stability scoring of stationary wavelet transform levels:

[0099]

[0100] in, Indicates the first The scoring results of the level of stationary wavelet transform. For subband stability, specifically, the median absolute deviation (MDR) is used. As a metric for the time-domain stability of each sub-band, The smaller the value, the higher the stability of the subband.

[0101] The stationary wavelet transform levels are sorted according to the stability score. A preset number of stationary wavelet transform levels are selected from the sorting results and inverse stationary wavelet transform is performed to obtain the LESR enhanced signal.

[0102] This invention embodiment is based on a score for all The levels are sorted using (stationary wavelet transform) and a fixed number of top-level levels are selected based on merit. In cases of identical scores, levels with lower frequencies are prioritized. Levels not selected are... After setting the coefficients to zero, an inverse stationary wavelet transform is performed to reconstruct the enhanced signal. This process effectively preserves the frequency band energy that is related to the click event and has stable time-domain characteristics, while significantly suppressing frequency band components that are severely affected by noise.

[0103] In another embodiment, to ensure the accuracy of biometric authentication, the present invention also provides a normalized segmentation method for signal segments, used to accurately separate click event segments from a continuous audio stream. The core of the method in this embodiment lies in utilizing the difference in sound field asymmetry between the tapping event and environmental noise, specifically including:

[0104] By acquiring the sound field asymmetry difference between the tapping events during user authentication and the ambient noise signal, the overall... parameter for:

[0105]

[0106] in, for The full-band energy of the top microphone channel at any given moment; for Full-band energy of the bottom microphone channel at all times; Used to characterize the degree of signal asymmetry between the channels of the top microphone and the channels of the bottom microphone;

[0107] Calculate the whole parameter The sample mean is used to select candidate tapping event frames from the sample mean by using a preset judgment threshold.

[0108] In some embodiments, the overall value within the analysis window is first calculated. sample mean Then, the judgment threshold was set as follows: ,in These are adjustable parameters calibrated on the validation set; ultimately, they will meet the conditions. The frames are identified as candidate tapping event frames. This method, which dynamically adjusts the threshold based on local signal characteristics, effectively improves the robustness and accuracy of the segmentation operation under different environmental noise conditions.

[0109] It is understood that the LESR enhanced signal and candidate knock event frames obtained in the embodiments of the present invention can both be used as noise-resistant biometrics.

[0110] The S300 uses pre-trained audio neural networks (PANNs) with enhanced sound field LESR to extract features from anti-noise biometrics, resulting in acoustic feature extraction results.

[0111] To achieve robust user authentication based on deep learning, the core challenge lies in learning an embedded representation that can effectively preserve identity-related information (such as differences between users) while suppressing irrelevant interference factors (such as noise, gestures, tapping force, and environmental changes) under a strict constraint of low false positive rate. Existing technical solutions have significant limitations: while using raw waveform encoders can capture fine-grained acoustic details, they are extremely sensitive to noise; relying on manual statistical features, although robust to some extent, lacks sufficient discriminative ability.

[0112] The key finding of this invention is that the sound field LESR model enables the fusion of the two types of features into a complementary representation. Based on this, embodiments of this invention propose a novel embedding learning framework with the following characteristics:

[0113] (1) Integrate complementary representations from pre-trained audio neural networks (PANNs) and sound field LESR (ΔLESR) to balance the discriminative power and robustness of features;

[0114] (2) The contrastive learning strategy is used to optimize the embedding representation, thereby effectively decoupling user identity features from interference factors.

[0115] In some embodiments, a LESR-enhanced PANNs feature extraction method is used. PANNs models were originally designed for general audio event recognition, capable of directly extracting high-dimensional time-frequency features from raw waveforms, thereby capturing subtle acoustic differences. In this invention, the identity information embedded in the click signal is considered a special kind of "acoustic content," making PANNs well-suited for distinguishing fine-grained biometric features generated by a user's finger taps. However, PANNs primarily focus on acoustic details and do not explicitly suppress interfering factors, thus they are relatively sensitive to noise and environmental changes. In contrast, It emphasizes session-level dynamics, capturing inherent biometrics and is robust to noise, finger posture, and typing force, thus highlighting the user's unique stability. However, if only using... However, this lacks sufficient fine-grained identity recognition capabilities. To combine the advantages of both, this invention integrates the fine-grained acoustic features extracted by PANNs with... The provided noise-resistant, behavior-independent, and structure-propagation-enhanced biometrics are spliced ​​together and processed through a lightweight fusion network consisting of two linear layers, a normalization layer, and a nonlinear activation function. This results in a fused embedding. This information is then normalized to a unit hypersphere. This design ensures that PANNs provide sensitive, microscopic identity information, while This anchors the stability of features under different sessions and environmental changes.

[0116] In some embodiments, the LESR-enhanced PANNs feature extraction method is as follows:

[0117] The PIN code click sound samples were extracted using pre-trained audio neural networks (PANNs) to obtain high-frequency time-frequency features;

[0118] High-frequency time-frequency features and noise-resistant biometric features are processed using a lightweight fusion network to obtain a fused embedding;

[0119] The fusion embedding is decoupled using contrastive learning to obtain acoustic feature extraction results, where the total loss function of contrastive learning is... for:

[0120]

[0121] The interval-based triplet loss function, for The weight parameters, The calculation method is as follows:

[0122]

[0123] in, For integration and embedding, This represents anchor samples, positive samples, and negative samples. Positive samples... With anchor point sample The same users, negative samples With anchor point sample Different users;

[0124] The binary cross-entropy loss function is... for The weight parameters, The calculation method is as follows

[0125]

[0126] in, It is the logit value predicted by the model. It's a label indicating whether the current sample is valid. These are the weight parameters.

[0127] This joint optimization method in this invention utilizes triplet loss for contrastive learning to maintain identity distinguishability, while simultaneously employing binary cross-entropy loss to ensure that the embedded representation is aligned with the final authentication target at the task level. These two methods work synergistically to decouple biometric features from interfering factors and improve the model's robustness to user behavior and environmental changes.

[0128] S400 uses single-sample classification and multi-sample authentication methods to classify the acoustic feature extraction results and obtain the PIN code identity authentication result.

[0129] The acoustic feature extraction results are converted into confidence scores using the sigmoid function. Among them, the confidence score This is used to characterize the probability that the PIN code click sound sample comes from the target user.

[0130] The authentication method implemented in this invention includes providing a sample for each click. Generate the corresponding embedding representation The classifier output is then converted into a calibrated confidence score using the sigmoid function. Used to characterize samples The probability that it originates from the target user.

[0131] In some embodiments, single-sample classification is based on confidence scores. The system uses a preset threshold to determine whether a single tap in the PIN code click sound sample comes from a legitimate user.

[0132] Understandably, for single-click authentication, if its confidence score... Not lower than the preset threshold If so, the tap is determined to have come from a legitimate user. The threshold is... Tune on the validation set to meet the system's intended operating point (e.g., maximizing the F1 score, equal error rate, or forcibly limiting the false positive rate). Intuitively, increase the threshold. This will reduce the risk of false acceptance, but may increase the probability of false rejection. During the inference phase, single-knock authentication is simplified to... Threshold-based decision-making is performed. Although this framework can effectively decouple identity-related features from changes caused by behavior and environment, single-sample authentication may still suffer from accidental misclassification due to the inherent probabilistic nature of deep learning models.

[0133] Multi-sample authentication based on confidence score The session-level acceptance probability is used to determine whether multiple PIN code click sounds originate from a legitimate user. This session-level acceptance probability includes the legitimate user session acceptance probability and the attacker session acceptance probability. The legitimate user session acceptance probability is:

[0134]

[0135] The attacker's session acceptance probability is:

[0136]

[0137] in, The number of keystrokes required for authentication. The number of times the PIN code was tapped to produce the sound sample. For the single-sample authentication success rate of legitimate users, The attacker's single-sample false positive rate. This is used as a sample serial number identifier.

[0138] In some embodiments, to improve authentication robustness, aggregation is performed within a single session. The click samples are used for joint decision-making, and each click is modeled as an independent Bernoulli trial.

[0139] refer to Figure 2 The basic architecture diagram of the identity authentication framework shows that when a mobile phone user clicks on the PIN code, this embodiment of the invention uses the PIN code and the sound of clicking the PIN code. As input to the authentication model. Firstly, this is proposed through an embodiment of the present invention. The model denoises the click sounds and obtains a complete dataset of all sub-bands and frames. Following the features, embodiments of the present invention construct a Feature matrix:

[0140]

[0141] The click sounds are then segmented using the adaptive threshold method proposed in this embodiment of the invention, thereby extracting the click sound for each PIN code. The audio signal of the click sounds is then processed... Figure 3 The diagram shown illustrates a PANNS-based feature extraction model used to extract an embedded feature, and... Together, they are processed through a lightweight fusion network consisting of two linear layers, a normalization layer, and a nonlinear activation function, resulting in a fusion embedding. . After passing through a fully connected layer and the sigmoid activation function The confidence score is used, and the final authentication result is calculated using the session-level acceptance probability.

[0142] While using the sound of clicking the PIN code for authentication, this embodiment of the invention also provides the original PIN code as a fallback. As long as the correct PIN code sequence is entered during authentication, authentication will pass even if the sound authentication fails, thereby improving the usability of the entire authentication system.

[0143] As shown in the experiment below, a large number of samples were collected from multiple participants using multiple devices in different environments. Each participant contributed two types of data: controlled input (at least 500 clicks on different PINs) and uncontrolled input across four environments (shopping mall, office, subway, taxi). The signals were recorded at 44.1 kHz using built-in dual microphones.

[0144] The evaluation metrics include True Positive Examples (TP), True Negative Examples (TN), False Positive Examples (FP), and False Negative Examples (FN). Unless otherwise stated, this embodiment uses the following four core evaluation metrics to quantitatively analyze system performance:

[0145] (1) Authentication Success Rate (ASR): refers to the proportion of legitimate users whose authentication attempts are successfully accepted. The calculation formula is as follows:

[0146]

[0147] (2) False Positive Rate (FPR): This refers to the proportion of attacks that are incorrectly accepted when an attacker attempts to impersonate a legitimate user. The formula is as follows:

[0148]

[0149] (3) F1 score: The harmonic mean of precision and recall, which comprehensively reflects the system’s ability to balance security (low FPR) and availability (low FNR).

[0150] (4) Equal Error Rate (EER): This refers to the error rate at which the false recognition rate (FPR) and the false rejection rate (FNR) of the system are equal at the operation point. It is a standard evaluation index widely used in the field of biometric identification and is used to comprehensively measure the overall verification accuracy of the system. The lower the EER value, the better the system performs in terms of balancing security and usability.

[0151] The above four indicators together constitute a comprehensive evaluation system for the performance of the identity authentication system proposed in this invention: ASR and FPR respectively characterize the system's usability and security; F1 score and EER summarize the system's comprehensive authentication performance in practical application scenarios from different dimensions.

[0152] The experimental results are shown below:

[0153] Figure 4 The experiment evaluated the performance of various metrics under different values ​​of N and M, iterating through the possible values ​​of M for different N values ​​(N∈{3, 5, 7, 10}), and conducted on a dataset of six test users. The results show that there is a broad near-saturation region: when N=5 and M=3, the system can achieve a high probability of legitimate user identification. Meanwhile, the attacker's success rate in impersonating others is high. Authentication can be completed with just five clicks. Increasing the M value enhances the system's resistance to interference, but it will slightly reduce the success rate for legitimate users. Decreasing the M value has the opposite effect. For the consecutive click variant (N=5), as the consecutive click length increases from 1 to 5, the pass rate for legitimate users decreases. The success rate dropped from 0.90 to 0.73, while the attacker's success rate remained relatively stable at approximately [missing value]. Scale. Based on the above experimental results, this invention uses a parameter combination of N=5 and M=3 in the default configuration to achieve the optimal balance between authentication efficiency, user-friendliness, and security.

[0154] Figure 5 The results show the recognition success rate and 1-false recognition rate for different users. It can be seen that each user has a high recognition success rate and a low false recognition rate, which demonstrates the effectiveness of the authentication method proposed in this invention.

[0155] Figure 6 The experimental results under common centralized scenarios demonstrate that the present invention has a certain degree of robustness to different scenarios and can be applied to common certification scenarios.

[0156] Figure 7 The results show the identification effectiveness at different times after user registration. It can be seen that even after three months, the authentication success rate is still very good and the false recognition rate is low, indicating the persistence of the identity-related features contained in the user's clicks.

[0157] Figure 8 This is a schematic diagram of a PIN code authentication system based on acoustic structural response according to an embodiment of the present invention. The system includes a first module 810, a second module 820, a third module 830, and a fourth module 840.

[0158] The system comprises four modules: the first module collects PIN code click sound samples during user authentication; the second module performs fine-grained denoising on the PIN code click sound samples to obtain noise-resistant biometric features, including key frequency band selection and normalized segmentation; the third module uses pre-trained audio neural networks (PANNs) with enhanced sound field LESR to extract features from the noise-resistant biometric features, resulting in acoustic feature extraction; and the fourth module classifies the acoustic feature extraction results using single-sample classification and multi-sample authentication methods to obtain the PIN code authentication result.

[0159] For example, with the cooperation of the first, second, third, and fourth modules in the system, the embodiment device can implement any of the aforementioned PIN code authentication methods based on acoustic structure response, namely, collecting PIN code click sound samples when a user performs identity verification; performing fine-grained denoising on the PIN code click sound samples to obtain noise-resistant biometric features, wherein fine-grained denoising includes key frequency band selection and normalized segmentation; using pre-trained audio neural networks (PANNs) with enhanced sound field LESR to extract features from the noise-resistant biometric features to obtain acoustic feature extraction results; and using single-sample classification and multi-sample authentication methods to classify the acoustic feature extraction results to obtain PIN code authentication results. This invention improves the security of users unlocking mobile terminals using PIN codes.

[0160] This invention also provides an electronic device, which includes a processor and a memory;

[0161] The memory stores the program;

[0162] The processor executes a program to perform the aforementioned PIN code authentication method based on acoustic structure response; the electronic device has the function of carrying and running the software system for PIN code authentication based on acoustic structure response provided in the embodiments of the present invention, such as a personal computer, minicomputer, mainframe, workstation, network or distributed computing environment, standalone or integrated computer platform, or communicating with charged particle tools or other imaging devices, etc.

[0163] This invention also provides a computer-readable storage medium storing a program that is executed by a processor to implement the PIN code authentication method based on acoustic structure response as described above.

[0164] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this invention are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented in the embodiments of this invention. Alternative embodiments are contemplated, in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0165] This invention also discloses a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium and execute the computer instructions, causing the computer device to perform the aforementioned PIN code authentication method based on acoustic structure response.

[0166] Furthermore, although the invention has been described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the invention. Rather, considering the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed in the embodiments of the invention, the actual implementation of the module will be understood within the scope of conventional skill of an engineer. Therefore, those skilled in the art can implement the invention as set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the invention, which is determined by the full scope of the appended claims and their equivalents.

[0167] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0168] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can include, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0169] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0170] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0171] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0172] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

[0173] The above is a detailed description of the preferred embodiments of the present invention, but the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A PIN code authentication method based on acoustic structural response, characterized in that, include: Collect sound samples of PIN code clicks during user authentication; Fine-grained denoising is performed on the PIN code click sound sample to obtain noise-resistant biometric features, wherein fine-grained denoising includes key frequency band selection and normalized segmentation; Pre-trained audio neural networks (PANNs) with enhanced sound field LESR were used to extract the noise-resistant biofeedback features, resulting in acoustic feature extraction. The acoustic feature extraction results are classified using single-sample classification and multi-sample authentication methods to obtain PIN code authentication results.

2. The PIN code authentication method based on acoustic structural response according to claim 1, characterized in that, The collected PIN code click sound samples during user authentication include: Obtain PIN code click sound samples captured by the top and bottom microphones of the mobile device: in The path-dependent channel frequency response; Represents frequency Click spectrum at the location; Represents environmental noise signals; for: in, Indicates the distance from the clicked location to the microphone. The effective transmission distance; Indicates geometric diffusion; This indicates the structural damping and scattering within the fuselage; This represents the finger-surface contact impedance and receiver chain gain.

3. The PIN code authentication method based on acoustic structural response according to claim 2, characterized in that, The method further includes: Click the PIN code to analyze the frequency band of the sound sample. The short-time Fourier transform is used to uniformly divide the area into... A continuous sub-band ; Based on the sub-band responses from the top and bottom microphones, determine the PIN code click sound sample in the sub-band. and the LESR features of frames for: in, For the child response of the top microphone; For the sub-band response of the bottom microphone; and To fix the click position, and This is a band-limited parameter.

4. The PIN code authentication method based on acoustic structural response according to claim 3, characterized in that, The selection of the key frequency band includes: Perform L-level stationary wavelet transform on the PIN code click sound samples from the top and bottom microphones to obtain frequency band components at different stationary wavelet transform levels, and then divide the subbands... By mapping the center frequency to the corresponding stationary wavelet transform level, a complete LESR feature matrix with sub-bands and frames is obtained, containing a set of mapping relationships. for: in, This serves as an identifier for the stationary wavelet transform level. For children The center frequency; The passband corresponds to the stationary wavelet transform level; the complete LESR characteristic matrix is: Stability scoring of stationary wavelet transform levels: in, Indicates the first The scoring results of the level of stationary wavelet transform. For subband stability; The stationary wavelet transform levels are sorted according to the stability score. A preset number of stationary wavelet transform levels are selected from the sorting results and inverse stationary wavelet transform is performed to obtain the LESR enhanced signal.

5. The PIN code authentication method based on acoustic structural response according to claim 3, characterized in that, The normalized segmentation includes: By acquiring the sound field asymmetry difference between the tapping events during user authentication and the ambient noise signal, the overall... parameter for: in, for The full-band energy of the top microphone channel at any given moment; for Full-band energy of the bottom microphone channel at all times; Used to characterize the degree of signal asymmetry between the channels of the top microphone and the channels of the bottom microphone; Calculate the whole parameter The sample mean is used to select candidate tapping event frames from the sample mean by using a preset judgment threshold.

6. The PIN code authentication method based on acoustic structural response according to claim 1, characterized in that, The pre-trained audio neural networks (PANNs) enhanced with LESR (Low Sound Field Resonance) are used to extract features from the noise-resistant biometrics, resulting in acoustic feature extraction results, including: The PIN code click sound samples were extracted using pre-trained audio neural networks (PANNs) to obtain high-frequency time-frequency features; High-frequency time-frequency features and noise-resistant biometric features are processed using a lightweight fusion network to obtain a fused embedding; The fusion embedding is decoupled using contrastive learning to obtain acoustic feature extraction results, where the total loss function of contrastive learning is... for The interval-based triplet loss function, for The weight parameters, The calculation method is as follows: in, For integration and embedding, This represents anchor samples, positive samples, and negative samples. Positive samples... With anchor point sample The same users, negative samples With anchor point sample Different users; The binary cross-entropy loss function is... for The weight parameters, The calculation method is as follows in, It is the logit value predicted by the model. It's a label indicating whether the current sample is valid. These are the weight parameters.

7. The PIN code authentication method based on acoustic structural response according to claim 6, characterized in that, The acoustic feature extraction results are classified using a single-sample classification and multi-sample authentication method to obtain PIN code authentication results, including: The acoustic feature extraction results are converted into confidence scores using the sigmoid function. Among them, the confidence score Used to characterize the probability that the PIN code click sound sample comes from the target user; The single-sample classification is based on confidence scores. Use preset thresholds to determine whether a single tap in the PIN code click sound sample comes from a legitimate user; The multi-sample authentication is based on confidence scores. The session-level acceptance probability is used to determine whether multiple PIN code click sounds originate from a legitimate user. This session-level acceptance probability includes the legitimate user session acceptance probability and the attacker session acceptance probability. The legitimate user session acceptance probability is: The attacker's session acceptance probability is: in, The number of keystrokes required for authentication. The number of times the PIN code was tapped to produce the sound sample. For the single-sample authentication success rate of legitimate users, The attacker's single-sample false positive rate. This is used as a sample serial number identifier.

8. The PIN code authentication method based on acoustic structural response according to claim 7, characterized in that, The preset threshold is determined based on the F1 score and error rate of the evaluation index.

9. A PIN code authentication system based on acoustic structural response, characterized in that, include: The first module is used to collect sound samples of PIN code clicks when users authenticate their identity. The second module is used to perform fine-grained denoising on the PIN code click sound sample to obtain noise-resistant biometric features, wherein fine-grained denoising includes key frequency band selection and normalized segmentation. The third module is used to extract features from the noise-resistant biofeedback using pre-trained audio neural networks (PANNs) enhanced with sound field LESR, and obtain acoustic feature extraction results. The fourth module is used to classify the acoustic feature extraction results using single-sample classification and multi-sample authentication methods to obtain PIN code authentication results.