On-screen interaction position-independent continuous identity authentication method based on acoustic multipath characteristics

By using active acoustic perception and adversarial learning networks, hand multipath response features are extracted during on-screen interaction, solving the problem that existing identity authentication methods rely on hardware or scene limitations. This enables continuous identity authentication that is independent of the on-screen interaction location, improving the reliability of authentication and user experience.

CN121959536APending Publication Date: 2026-05-01BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF TECH
Filing Date
2026-01-23
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing authentication methods for mobile smart devices rely on specific hardware or scenarios, which limits their application scenarios, affects user experience, and makes it difficult to achieve simple and easy-to-use continuous authentication.

Method used

By using active acoustic sensing technology, acoustic signals are generated using ZC sequences. Combined with bandpass filtering and cross-correlation processing, hand multipath response features are extracted during on-screen interaction. Furthermore, an adversarial learning network is used to train a feature extractor and an identity recognizer to achieve position-independent identity authentication during on-screen interaction.

Benefits of technology

It enables continuous identity authentication across different screen positions and interaction trajectories, improving the reliability of identity authentication and user experience without relying on additional hardware devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121959536A_ABST
    Figure CN121959536A_ABST
Patent Text Reader

Abstract

The invention discloses an on-screen interaction position-independent continuous identity authentication method based on acoustic multipath characteristics, and belongs to the technical field of mobile intelligent terminal security and identity authentication. In order to solve the problem that an existing method depends on an additional sensor or is limited by a specific interaction scene, the method adopts a ZC sequence with high autocorrelation to modulate an active acoustic sensing signal, and through technologies such as band-pass filtering, direct propagation path suppression and cross-correlation, multipath response corresponding to hand postures is separated and extracted from collected sensing audio. According to the method, multipath responses at a plurality of on-screen interaction areas are used as training samples, an adversarial learning network composed of a feature extractor, an area recognizer and an identity recognizer is constructed, and joint training is carried out, so that the extracted identity features inhibit information related to on-screen interaction positions, and the identity recognition capability is improved at the same time; therefore, continuous identity authentication irrelevant to the interaction position on the screen is realized. The method has the effects of being irrelevant to on-screen interaction positions, anti-noise, adaptive to continuous authentication scenes and the like.
Need to check novelty before this filing date? Find Prior Art

Description

A screen-based interactive location-independent persistent authentication method based on acoustic multipath features Technical Field

[0001] This invention belongs to the field of mobile intelligent terminal security and identity authentication, and relates to an on-screen interactive location-independent continuous identity authentication method based on acoustic multipath features, which combines acoustic perception and adversarial learning networks. It is applicable to secure access authentication scenarios for various terminal devices. Background Technology

[0002] To prevent the misuse of mobile smart devices and the resulting leakage of users' personal privacy and economic losses, mobile smart devices generally support various new authentication methods. For example, instant authentication methods based on facial structure, mouth movements, heartbeat, and finger internal structures generally require active user cooperation to complete the authentication process, thus affecting the user experience. Compared to methods for instant authentication, authentication methods based on heartbeat and gait characteristics for continuous authentication can meet more complex user needs. However, these methods rely on additional hardware devices (such as wristband PPG devices) or can only be applied in specific scenarios (such as walking). Furthermore, during on-screen interaction, the shape of users' hand gestures varies individually, thus serving as a biometric feature for user identification. Meanwhile, mobile smart devices are generally equipped with speakers and microphones, enabling them to perceive environmental spatial features through active audio perception. Therefore, existing research has utilized active audio perception technology to obtain spatially specific biometric features related to hand gestures during on-screen interaction, achieving implicit authentication for mobile devices. However, these studies are limited by specific interaction locations on the screen, resulting in limited application scenarios. Therefore, how to achieve simple, easy-to-use, non-specific sensor-required, and scenario-independent continuous identity authentication based on commercial smart devices is a crucial issue that needs to be addressed to ensure the security of mobile terminals. Summary of the Invention

[0003] This invention addresses the shortcomings of existing hand gesture authentication methods for on-screen interactions in smart devices. It extracts fine-grained acoustic multipath response features related to the hand during on-screen interactions based on highly autocorrelation ZC sequence sensing acoustic signals. Then, it achieves reliable and continuous authentication that is position-independent, requires no additional hardware, and utilizes adversarial learning.

[0004] The innovation of this invention lies in two aspects: First, the multipath reflection signals collected by active acoustic sensing include not only hand reflection signals but also reflection signals from the user's torso, pedestrians, and other facilities in the environment. Therefore, this invention designs a ZC sequence-based acoustic signal sensing system and uses a series of signal processing techniques to separate the multipath reflection signals corresponding only to interactive gestures from the audio signals collected by active audio sensing. Second, during on-screen interaction, the user's hand is constantly moving, and the corresponding multipath reflection paths change dynamically accordingly. Therefore, this invention utilizes an adversarial learning network to train the acoustic multipath features of interactive gestures, extracting hand posture features independent of the on-screen interaction location. Ultimately, this achieves simple, easy-to-use, and continuous hand posture authentication that is not limited to the on-screen interaction location.

[0005] The objective of this invention is achieved through the following technical solutions.

[0006] Step 1: Acoustic Sensing Signal Generation. During the registration and authentication phases, this invention generates acoustic sensing signals using a baseband ZC sequence. The generation process includes generating the ZC sequence baseband signal, upsampling, and carrier modulation. This step yields a high signal-to-noise ratio, wide-bandwidth acoustic signal, providing a fundamental signal source for identity feature extraction during subsequent on-screen interactive sensing processes.

[0007] Step Two: On-Screen Interaction Sample Collection. During the registration phase, generated acoustic sensing signals are used to perceive user gestures at different interactive locations on the screen. The coordinates of each interactive location and the corresponding collected acoustic sensing samples are recorded, forming a registration sample set. During the authentication phase, acoustic sensing signals are also used to collect user interaction samples, providing input samples for identity recognition.

[0008] Step 3: Sensing Signal Preprocessing. Environmental noise cancellation and hand multipath feature extraction are performed on the collected acoustic samples. Out-of-band noise is eliminated through bandpass filtering to obtain clean audio segments; subsequently, direct propagation path cancellation and cross-correlation operations are used to extract multipath response features corresponding to interactive gestures, thereby highlighting signal features related to individual user differences and providing high-quality input information for identity feature extraction.

[0009] Step 4: Training the Interactive Location-Independent Authentication Model. The adversarial learning network used consists of three parts: a feature extractor, a region recognizer, and an identity recognizer. During the adversarial learning process, the feature extractor, region recognizer, and identity recognizer are jointly trained, while suppressing region-related features and improving identity recognition capabilities, so that the extracted identity features are independent of on-screen interactive location.

[0010] Step 5: User Identity Authentication. The acoustic sample provided by the user to be authenticated undergoes noise reduction and multipath feature extraction. It is then input into a pre-saved feature extractor to extract identity features, and finally into a saved identity recognizer for authentication. If the maximum predicted probability output by the recognizer is higher than a preset security threshold, the user corresponding to that maximum probability is identified as the user to be authenticated; otherwise, they are identified as a stranger, thus achieving secure and reliable user identity authentication.

[0011] Compared with the prior art, the present invention has the following advantages:

[0012] 1. By employing ZC sequence active acoustic sensing signals and combining bandpass filtering, direct path suppression, and cross-correlation processing, this invention can accurately extract fine-grained features of hand multipath response, providing rich and reliable discrimination information for identity recognition.

[0013] 2. This invention utilizes convolutional blocks and bidirectional LSTM feature extractors in conjunction with adversarial training of region and identity recognition units to suppress information related to interaction location and retain only individual user features, thereby achieving continuous identity authentication under different screen locations and interaction trajectories.

[0014] 3. This invention relies solely on the built-in speaker and microphone of the mobile terminal, requiring no external sensors. Combined with a highly efficient feature extraction network, it can implement continuous identity authentication in real-time or near-real-time scenarios, demonstrating good feasibility and engineering application value. Attached Figure Description

[0015] Figure 1 is a diagram illustrating the authentication method architecture of an embodiment of the present invention.

[0016] Figure 2 is a diagram of the adversarial learning network architecture according to an embodiment of the present invention.

[0017] Figure 3 shows the authentication accuracy of the embodiments of the present invention.

[0018] Figure 4 shows the false rejection rate of the embodiments of the present invention.

[0019] Figure 5 shows the error acceptance rate of the embodiments of the present invention. Detailed Implementation

[0020] As shown in Figure 1, a screen-based interactive location-independent persistent authentication method based on acoustic multipath features includes the following steps:

[0021] Step 1: Generation of Acoustic Sensing Signals

[0022] In the authentication method of this invention, both the registration and authentication phases involve the generation of acoustic sensing signals. Among various baseband signal types, ZC sequences have superior autocorrelation gain. Therefore, this invention selects ZC sequences as the baseband signal to modulate the acoustic sensing signal. The modulation of ZC sequence acoustic sensing signals mainly consists of three steps: ZC baseband signal generation, upsampling, and carrier modulation.

[0023] Step 1.1 Generation of ZC sequence baseband signal: length is ZC sequence Defined as .in, It is the imaginary unit. The definition is as follows:

[0024]

[0025] in, For modulo operation, and Two integers that are coprime, and .in, The number of chirp signals contained in the predefined ZC sequence passband signal. This invention employs... The sampling frequency of the sensing signal, and the starting frequency of the sensing signal. With termination frequency The frequencies were set to 17kHz and 23kHz respectively. This invention will increase the length of the sensed signal sample. Set to 2048 (approximately 43ms). Correspondingly, the ZC sequence... length =257. Number of chirps in the ZC sequence. In the case of 4, .

[0026] Step 1.2 Upsampling: In this step, the present invention uses frequency domain zero-padding to upsample the ZC sequence signal. Upsampling is then performed. Specifically, the generated ZC sequence is first subjected to a Fast Fourier Transform, which outputs frequency domain information of length 257. Then fill the frequency domain according to the following formula. A zero element, obtain .

[0027]

[0028] in, and They represent Extracting from the first element up to the th The fragment of the element and from the first element The element is started to be captured. The final segment. Indicates generation One zero element. Indicates will , and The three segments are pieced together in sequence.

[0029] Finally, through the analysis of Perform inverse Fourier transform to obtain the corresponding time-domain signal At this point, the upsampled ZC sequence length The value is 2048, and the frequency domain bandwidth is also limited to 6kHz.

[0030] Step 1.3 Carrier Modulation: This step modulates the upsampled ZC sequence time-domain signal with a bandwidth of 6kHz to 17kHz to 23kHz. First, a frequency of... I component of the carrier signal With Q component . The signal time variable and Then the real part of the upsampled ZC time-domain complex signal is... With the imaginary part The corresponding acoustic sensing signals can be obtained by calculating them using the following formulas:

[0031]

[0032] Step 2: On-screen interaction sample collection

[0033] The registration phase requires acoustic sensing of user gestures and postures at multiple interactive locations on the screen. This acoustic sensing is accomplished by the combination of the top speaker and bottom microphone of the mobile smart device. Specifically, during one acoustic sensing operation, the top speaker plays an acoustic sensing signal. The signal is recorded by the bottom microphone of the device through a direct propagation path, and also by the bottom microphone after reflection from the hand and the environment. During the registration sample collection process, users need to swipe in a serpentine motion from top to bottom on the mobile device screen at a fixed speed with a consistent hand posture, ensuring that the swipe trajectory covers the entire interactive area of ​​the screen. During this process, the first signal is captured using the smart device's touchscreen. Two interactive position coordinates (in and (representing the x and y coordinates of the on-screen interaction location and the corresponding acoustic perception samples) Furthermore, to achieve position-independent on-screen interaction, this invention divides the on-screen interaction area of ​​the mobile device into... There are a total of 9 areas. Based on the interaction location coordinates... Mapped to the corresponding interactive area label Specifically, let the current screen interaction area be a rectangular area, with horizontal and vertical dimensions of [missing information]. and Given coordinates Its corresponding interactive area index ,in This indicates taking the minimum value. This indicates rounding down. Then, the interactive area label... With the corresponding acoustic perception samples Composition of the registration sample set This is used to complete subsequent user registration.

[0034] Acoustic sensing signals are also required during the authentication phase. Acoustic perception samples from on-screen interactions are collected for user identification. Unlike the registration phase, the authentication phase does not require collecting the corresponding on-screen interaction location coordinates.

[0035] Step 3: Sensing signal preprocessing

[0036] This step mainly involves removing environmental noise and extracting hand multipath features from the acoustic perception samples collected during the registration and authentication phases.

[0037] Step 3.1 Ambient noise cancellation: This invention utilizes a Butterworth bandpass filter. Based on the preset starting frequency of the acoustic sensing signal in step one With termination frequency Through acoustic perception samples Out-of-band noise cancellation removes environmental noise interference, thereby obtaining a clean audio clip. .Right now:

[0038]

[0039] Step 3.2 Multipath Feature Extraction: This step first uses spectral subtraction to eliminate components of the direct propagation path during the acoustic perception of hand gestures, thereby improving the salience of personalized multipath features corresponding to hand gestures. If... If we represent the acoustic segment after eliminating the direct transmission path, then:

[0040]

[0041] middle, and These represent the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively. The acoustic sensing signal generated in step one.

[0042] Then, the present invention utilizes cross-correlation operations. right Processing is performed to obtain the corresponding multipath effect response. .Right now:

[0043]

[0044] in, This indicates taking the absolute value. On one hand, the output of the cross-correlation operation is symmetrical about the center point. On the other hand, it's necessary to exclude multipath responses corresponding to the environment from the cross-correlation output, thus retaining only the multipath responses related to hand reflexes. Therefore, this step... Starting from the center point, 128 samples corresponding to the fragments are extracted to the right and used as the response signal containing only hand reflection multipath.

[0045] Step 4: Training the Interaction Location-Independent Authentication Model

[0046] During the registration phase, this step trains the hand reflection multipath response obtained from the above steps on the registration samples using adversarial learning, thereby achieving an on-screen interaction position-independent authentication model. As shown in Figure 2, the adversarial learning network used in this invention includes three parts: a feature extractor, a region identifier, and an identity recognizer. During the adversarial learning process, the feature extractor, region identifier, and identity recognizer are jointly trained and their parameters are updated synchronously. After training, the corresponding user-personalized feature extractor and identity recognizer are saved in the user authentication model database for use in the authentication phase to verify the identity of the user to be authenticated, thus achieving reliable user identity authentication. The specific design is as follows:

[0047] Step 4.1 Feature Extractor: To effectively extract fine-grained acoustic multipath features and temporal contextual information related to user interaction gestures, the feature extractor used in this invention consists of three convolutional blocks and two BiLSTM layers. Each convolutional block comprises a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a pooling layer. The output channels of the three convolutional modules are 64, 128, and 256, respectively, with a kernel size of 3 and a stride of 1. The pooling layer has a window size of 2 and a stride of 2. The two BiLSTM layers have 256 and 128 hidden units, respectively.

[0048] Step 4.2 Region Recognizer: The region recognizer consists of two fully connected layers and one softmax layer. The first fully connected layer contains 256 neurons. Since the interactive area on the screen is divided into 9 regions during the registration phase, the second fully connected layer contains 9 neurons. The relevant loss function is defined using cross-entropy:

[0049]

[0050] in, The first one represented by one-hot encoding The true probability distribution of each region on the screen (1 indicates that it belongs to the region, and 0 indicates that it does not). For the region identifier to the first The predicted probability of each area on the screen. The preset number of screen-mounted nodes is 9 here. Step 4.3 Identity Recognizer: The identity recognizer has the same structure as the region recognizer, consisting of two fully connected layers and one softmax layer. The first fully connected layer also includes 256 neurons. The number of neurons in the second fully connected layer is consistent with the number of users participating in training. Here, it is set to 10. The relevant loss function is defined using cross-entropy:

[0051]

[0052] in, The first one represented by one-hot encoding The true probability distribution of each user (1 indicates that the user is identified, and 0 indicates that the user is not identified). For the identity recognition device to the first The predicted probability for each user. The number of users participating in the training (10 in this case).

[0053] Step 4.4 Adversarial Network Training. The goal of adversarial learning training is to improve the user identity feature extraction capability by maximizing the accuracy of the identity recognizer, while suppressing the recognition probability of the region recognizer to achieve on-screen interaction location-independent identity authentication. Therefore, given the region recognizer loss function ( ) and the identity recognition loss function ( After that, the overall loss function for adversarial learning is:

[0054]

[0055] in, and Empirically set to 0.8 and 0.2. Left-handed, by minimizing... This improves identity recognition performance and suppresses region-related features.

[0056] Step 5: User Identity Verification

[0057] During the authentication phase, the acoustic perception samples provided by the user to be authenticated first undergo environmental noise cancellation and multipath feature extraction processing in step three to obtain the corresponding multipath response signal. The multipath response is then input into a feature extractor stored in the registration phase for feature encoding to extract a discriminative user identity feature representation. Subsequently, this identity feature representation is input into an identity recognizer stored in the registration phase to complete user identity verification. Specifically, when the maximum predicted probability in the softmax output of the identity recognizer is greater than a preset security threshold of 0.9, the user identity corresponding to this maximum predicted probability is determined to be the user to be authenticated; when the maximum predicted probability is not greater than the security threshold, the user to be authenticated is determined to be a stranger, thereby improving the system's security and anti-impersonation capabilities.

[0058] Example

[0059] This embodiment invited 12 volunteers (7 men and 5 women) to collect experimental samples in a laboratory and a conference room using commercial smartphones: a REDMI K40 (hereinafter referred to as Device 1) and an ONEPLUS Ace5 (hereinafter referred to as Device 2). During sample collection, each volunteer chose a comfortable hand posture and swiped on the phone screen, ensuring the swipe trajectory covered the entire interactive area. While swiping, volunteers were required to maintain a consistent hand posture. The corresponding touch position coordinates and acoustic perception samples were then collected to form the experimental sample set (more than 8400 valid samples were collected). During the experimental verification process, two volunteers were selected as strangers, and the remaining 10 volunteers were selected as legitimate users. Then, up to 80% of the samples from each legitimate user's screen area were used for registration, and the remaining 20% ​​were used for authentication testing.

[0060] Figure 3 shows the average authentication accuracy in laboratory and conference room scenarios as the sample size changes during the legitimate user registration phase (covering both correct identification of legitimate users and correct identification of strangers). With the increase in the number of participating registration samples, the average authentication accuracy in both laboratory and conference room scenarios shows a continuous upward trend. When the registration sample ratio is low (20%), the accuracy in the conference room scenario is higher than that in the laboratory scenario. However, when the registration sample ratio is high (80%), the accuracy in the laboratory scenario (approximately 0.94) approaches that of the conference room scenario (approximately 0.95). This result indicates that increasing the registration sample ratio can effectively improve authentication performance and also verifies the reliability of the authentication method designed in this invention.

[0061] Figure 4 shows the false rejection rate (the probability that a legitimate user is identified as a stranger) when 80% of the samples are used for registration by legitimate users. The figure shows that in both the conference room and laboratory scenarios, using device 1 and device 2, the false rejection rate is less than 1%, indicating that the authentication method of the present invention can ensure a good user experience for registered users when users provide a certain number of training samples.

[0062] Figure 5 shows the false acceptance rate (the probability of a stranger being authenticated as a legitimate user) when 80% of the samples are registered by legitimate users. The figure shows that in both conference room and laboratory scenarios, using device 1 and device 2, the false rejection rate is less than 0.8%, indicating that the authentication method of this invention can effectively prevent unfamiliar users from passing authentication.

Claims

1. A method for on-screen interactive location-independent persistent authentication based on acoustic multipath features, characterized in that, Includes the following steps: Step Step 1: Acoustic sensing signal generation; During the registration and authentication phases, acoustic sensing signals are generated using baseband ZC sequences; The generation process includes the generation, upsampling, and carrier modulation of the ZC sequence baseband signal; Through this step, high signal-to-noise ratio and wide bandwidth acoustic signals can be obtained, providing a basic signal source for identity feature extraction in the subsequent on-screen interaction sensing process; Step 2: On-screen interaction sample collection; During the registration phase, the generated acoustic sensing signals are used to sense the user's gestures at different interactive positions on the screen, and the coordinates of each interactive position and the corresponding collected acoustic sensing samples are recorded to form a registration sample set. During the authentication phase, acoustic sensing signals are also used to collect user interaction samples to provide input samples for identity recognition; Step 3: Sensing signal preprocessing; Environmental noise cancellation and hand multipath feature extraction are performed on the collected acoustic samples; By eliminating out-of-band noise through bandpass filtering, a clean audio clip can be obtained; Subsequently, direct propagation path elimination and cross-correlation operations are used to extract multipath response features corresponding to interactive gestures, thereby highlighting signal features related to individual user differences and providing high-quality input information for identity feature extraction; Step 4: Training of the interaction location-independent authentication model; The adversarial learning network used includes three parts: a feature extractor, a region recognizer, and an identity recognizer; During the adversarial learning process, the feature extractor, region recognizer, and identity recognizer are jointly trained, while suppressing region-related features and improving identity recognition capabilities, so that the extracted identity features are independent of the on-screen interaction location; Step 5: User identity authentication; After noise cancellation and multipath feature extraction, the acoustic samples provided by the user to be authenticated are input into the pre-saved feature extractor to extract identity features, and then input into the saved identity recognizer for identity authentication; If the maximum predicted probability output by the identifier is higher than the preset security threshold, then the user corresponding to the maximum probability is determined to be the user identity to be authenticated. Otherwise, the user is identified as a stranger, thus achieving secure and reliable user authentication.

2. The method according to claim 1, characterized in that: Step 1: Acoustic Sensing Signal Generation Step 1.1 Generation of ZC Sequence Baseband Signal: Length is ZC sequence Defined as ;in, It is the imaginary unit; The definition is as follows: in, For modulo operation, and Two integers that are coprime, and ;in, The number of chirp signals contained in the predefined passband signal of the ZC sequence; ZC sequence length =257; Step 1.2 Upsampling: In this step, the ZC sequence signal is upsampled using frequency-domain zero-padding. Upsampling is performed; specifically, the generated ZC sequence is first subjected to a Fast Fourier Transform to output frequency domain information. Then fill the frequency domain according to the following formula. A zero element, obtain ; in, and They represent Extracting from the first element up to the th The fragment of the element and from the first element The element is started to be captured. The final segment; Indicates generation One zero element; Indicates will 、 and The three segments are pieced together in sequence; finally, through the analysis of... Perform inverse Fourier transform to obtain the corresponding time-domain signal Step 1.3 Carrier Modulation: Generate frequency of I component of the carrier signal With Q component ; The signal time variable and Then the real part of the upsampled ZC time-domain complex signal is... With the imaginary part The corresponding acoustic sensing signals can be obtained by calculating them using the following formulas:

3. The method according to claim 1, characterized in that: Step Two: The on-screen interaction sample collection and registration phase requires acoustic sensing of user gestures and postures at multiple interactive locations on the screen. This acoustic sensing is accomplished by the combination of the top speaker and bottom microphone of the mobile smart device. Specifically, during one acoustic sensing process, the top speaker of the device plays an acoustic sensing signal. The signal is recorded by the microphone at the bottom of the device through a direct propagation path, and also by the microphone at the bottom after being reflected by the hand and the environment. During the registration sample collection process, users need to use a consistent hand gesture to swipe in a serpentine motion from top to bottom on the mobile device screen at a fixed speed, ensuring the swipe trajectory covers the entire interactive area of ​​the screen. During this process, the touchscreen of the smart device is used to acquire the first... Two interactive position coordinates in and These represent the x and y coordinates of the on-screen interaction location, respectively; and the corresponding acoustic perception samples. Furthermore, to achieve position-independent on-screen interaction, the on-screen interaction area of ​​the mobile device is divided into... There are a total of 9 areas; based on the interaction location coordinates Mapped to the corresponding interactive area label Specifically, let the current screen interaction area be a rectangular area, with horizontal and vertical dimensions of respectively. and Given coordinates Its corresponding interactive area index ,in This indicates taking the minimum value; This indicates rounding down; then, the interactive area label... With the corresponding acoustic perception samples Composition of the registration sample set This is used to complete subsequent user registration; acoustic sensing signals are also required during the authentication phase. Acoustic perception samples during on-screen interactions are collected for user identification; unlike the registration phase, the authentication phase does not require the collection of corresponding on-screen interaction location coordinate information.

4. The method according to claim 1, characterized in that: Step 3: Sensing Signal Preprocessing. This step mainly involves removing environmental noise and extracting hand multipath features from the acoustic sensing samples collected during the registration and authentication phases. Step 3.1 Ambient noise cancellation: using a Butterworth bandpass filter Based on the preset starting frequency of the acoustic sensing signal in step one With termination frequency Through acoustic perception samples Out-of-band noise cancellation removes environmental noise interference, thereby obtaining a clean audio clip. ;Right now: Step 3.2 Multipath Feature Extraction: If we represent the acoustic segment after eliminating the direct transmission path, then: in, and These represent the Fast Fourier Transform and the Inverse Fast Fourier Transform, respectively. The acoustic sensing signal generated in step one; then, cross-correlation operation is used. right Processing is performed to obtain the corresponding multipath effect response. ;Right now: in, This indicates taking the absolute value; on the one hand, the output of the cross-correlation operation is symmetrical about the center point; on the other hand, it is necessary to exclude the multipath response corresponding to the environment in the cross-correlation output, so as to retain only the multipath response related to hand reflexes; therefore, this step is based on... Starting from the center point, multiple segments corresponding to the samples are extracted to the right and used as the response signal containing only hand reflection multipath.

5. The method according to claim 1, characterized in that: Step 4: The specific design for training the interaction location-independent authentication model is as follows: Step 4.1 Feature Extractor: To effectively extract fine-grained acoustic multipath features and temporal contextual information related to user interaction gestures, the feature extractor consists of three convolutional blocks and two BiLSTM layers; each convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, a ReLU activation layer, and a pooling layer; the output channels of the three convolutional modules are 64, 128, and 256, respectively, the kernel space size is 3, and the stride is 1; the pooling layer window is set to 2, and the stride is 2; the hidden units of the two BiLSTM layers are 256 and 128, respectively; Step 4.2 Region Recognizer: The region recognizer consists of two fully connected layers and one softmax layer; the first fully connected layer includes 256 neurons; since the on-screen interaction area is divided into 9 regions during the registration phase, the second fully connected layer includes 9 neurons; the relevant loss function is defined by cross-entropy: in, The first one represented by one-hot encoding The true probability distribution of each region on the screen, where 1 indicates belonging to the region and 0 indicates not belonging to it; For the region identifier to the first Predicted probability of each area on the screen; The preset number of screen-mounted neurons; Step 4.3 Identity Recognizer: The identity recognizer has the same structure as the region recognizer, consisting of two fully connected layers and one softmax layer; the first fully connected layer also includes 256 neurons; the number of neurons in the second fully connected layer is the same as the number of users participating in training, set to 10; the relevant loss function is defined using cross-entropy: in, The first one represented by one-hot encoding The true probability distribution of each user, where 1 indicates that the user is identified and 0 indicates that the user is not identified; For the identity recognition device to the first Predicted probability for each user; The number of users participating in the training is 10; Step 4.4 Adversarial Network Training; The goal of adversarial learning training is to improve the user identity feature extraction capability by maximizing the accuracy of the identity recognizer, while suppressing the recognition probability of the region recognizer to achieve on-screen interaction location-independent identity authentication; Therefore, given the region recognizer loss function ( ) and the identity recognition loss function ( After that, the overall loss function for adversarial learning is: in, and Empirically, we set it to 0.8 and 0.

2.

6. The method according to claim 1, characterized in that: Step 5: User Identity Authentication. In the authentication phase, the acoustic perception samples provided by the user to be authenticated first undergo environmental noise cancellation and multipath feature extraction processing in Step 3 to obtain the corresponding multipath response signal. The multipath response is then input into the feature extractor stored in the registration phase for feature encoding to extract a distinctive user identity feature representation. Subsequently, the identity feature representation is input into the identity recognizer stored in the registration stage to complete user identity authentication. Specifically, when the maximum predicted probability in the softmax output of the identity recognizer is greater than the preset security threshold of 0.9, the user identity corresponding to the maximum predicted probability is determined to be the identity of the user to be authenticated. When the maximum predicted probability is not greater than the security threshold, the user to be authenticated is determined to be a stranger, thereby improving the security and anti-impersonation capabilities of the system.