An Implicit Authentication Method for Mobile Devices Based on Attention Mechanism
By adopting an implicit authentication method based on attention mechanism on mobile devices, combining data augmentation, dynamic differential attention mechanism and model distillation technology, the problems of noise interference and insufficient model performance in the existing technology are solved, and higher authentication accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510346011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing implicit authentication methods for mobile devices have shortcomings in noise interference and model performance control, and it is difficult to effectively reduce noise interference and improve the performance of the authentication model.
Using an implicit identity authentication method for mobile devices based on attention mechanism, the distillation position of the teacher model is dynamically selected to adapt to the equipment performance through multi-source data preprocessing, generation-identification framework data augmentation, sliding window processing and cross-attention mechanism fusion characteristics, combined with dynamic differential attention mechanism and model distillation technology.
It significantly reduces noise interference, improves the recognition accuracy and generalization ability of the authentication model, moderately enhances the model's scenario coverage, accuracy and efficiency, and reduces the demand for device performance.
Smart Images

Figure CN119862554B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of network security, deep learning, and sensor technology, and specifically relates to an implicit authentication method for mobile devices based on an attention mechanism, which uses the data collected by sensors to perform implicit identity recognition and verification. Background Art
[0002] In recent years, with the research and progress of hardware, lightweight devices have become increasingly popular and play an irreplaceable role in life. Nowadays, the role of lightweight devices as a bridge between different users and data is becoming increasingly obvious, and they store a large amount of user privacy information, including mobile banking, communication, and sensitive personal data.
[0003] Identity authentication plays a crucial role in protecting the privacy of users in lightweight devices and is a key link to ensure the authenticity of user identities. The existing identity authentication mechanism is the entry method. When a user uses the device for the first time, through a verification strategy based on knowledge factors, the user is required to remember and enter specific information (for example, passwords, personal identification numbers PINs, and graphical gestures) or physiological biotechnology (for example, fingerprints, faces, irises, or voiceprints) for initial identity authentication.
[0004] However, research shows that the entry method is vulnerable to attacks (for example, brute-force attacks, shoulder-surfing attacks, touch-screen stains, and sensor-based judgments). Therefore, many existing methods use the environmental characteristics of the user using the device (for example, WIFI and GPS, or user behavior biometrics based on the motion sensors built into the mobile phone), which has led to extensive research on implicit user authentication systems that run continuously in the background. However, there are still some deficiencies in data noise processing and, at the same time, in controlling the performance of the model. Summary of the Invention
[0005] The purpose of the present invention is to provide an implicit authentication method for mobile devices based on an attention mechanism, which can effectively reduce noise interference and improve the performance of the model used for identity authentication.
[0006] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0007] An implicit authentication method for mobile devices based on an attention mechanism, the implicit authentication method for mobile devices based on an attention mechanism includes:
[0008] Collect multi-source data generated when a mobile device is used by multiple motion sensors of the mobile device and perform preprocessing;
[0009] Use a generative-discriminative framework to perform data augmentation on the preprocessed multi-source data to obtain multi-source augmented data;
[0010] Process multi-source enhanced data using a sliding window, and adopt a cross-attention mechanism to convert the multi-source enhanced data within a window into fused features, and use the labeled fused features to train a teacher model;
[0011] Obtain a student model by distilling the trained teacher model. During distillation, dynamically select the distillation position of the teacher model according to the performance score of the mobile device;
[0012] Input the fused features to be verified into the student model to obtain the authentication result output by the student model, and the authentication result is an illegal user or a legitimate user.
[0013] The following also provides several optional methods, which are not additional limitations to the above overall solution, but are only further supplements or optimizations. Without technical or logical contradictions, each optional method can be combined with the above overall solution alone, or multiple optional methods can be combined with each other.
[0014] Preferably, the generation-discrimination framework includes a generator, a transformer encoder, and a discriminator.
[0015] Preferably, the teacher model is a Transformer model, and the Transformer model adopts a dynamic differential attention mechanism. The attention calculation formula of the dynamic differential attention mechanism is:
[0016]
[0017] In the formula, is the attention value, is the query vector, is the key vector, is the value vector, is the softmax function, is the transpose of the key vector , is the dimension of the key vector , is the dynamic differential parameter, and the dynamic differential parameter is adjusted as follows:
[0018] Take the multi-source enhanced data within a window and calculate the variance of the enhanced data corresponding to each motion sensor;
[0019] If the variance of the enhanced data of a motion sensor is greater than the reasonable fluctuation range, increase the dynamic differential parameter ; otherwise, decrease the dynamic differential parameter , and the dynamic differential parameter .
[0020] Preferably, the method of generating the position embedding vector based on the rotary position encoding of the Transformer model is replaced by the method of generating the position embedding vector based on the learned position encoding. The calculation formula for generating the position embedding vector based on the learned position encoding is as follows:
[0021]
[0022] In the formula, is the position embedding vector, is the th time step, , is the total number of time steps, is the Embedding function.
[0023] Preferably, the calculation formula of the RMSNorm normalization mechanism of the Transformer model is as follows:
[0024]
[0025] In the formula, is the normalized eigenvalue output for the input feature , is the number of elements in the input feature , is the th element in the input feature , is a constant, is the scaling ratio, is the offset.
[0026] Preferably, obtaining the student model by distilling the trained teacher model includes:
[0027] Calculating the performance score of the mobile device;
[0028] Assume that the teacher model has a total of layers. If the performance score is less than the score threshold, take the first layers of the teacher model as the distillation range; otherwise, take the last layers of the teacher model as the distillation range;
[0029] For the distillation range, use the KMeans clustering method for clustering, and take the centroid of each cluster in the clustering result as the distillation position;
[0030] Train and output the student model based on the distillation position in the teacher model.
[0031] Preferably, the number of clusters of the KMeans clustering method is determined as follows:
[0032] Based on the maximum and minimum values of the performance scores, three score ranges are divided: a high score range, a medium score range, and a low score range;
[0033] According to the calculated performance score of the mobile device, determine the score range of the mobile device;
[0034] If the score range of the mobile device is the high score range, determine that the number of clusters is the first numerical value;
[0035] Alternatively, if the score range of the mobile device is the medium score range, determine that the number of clusters is the second numerical value;
[0036] Alternatively, if the score range of the mobile device is the low score range, determine that the number of clusters is the third numerical value; and the first numerical value is greater than the second numerical value, and the second numerical value is greater than the third numerical value.
[0037] Preferably, for the distillation range, the KMeans clustering method is used for clustering, including:
[0038] Randomly sample the layers within the distillation range;
[0039] Based on the random sampling results, use the KMeans clustering method for clustering.
[0040] An implicit authentication method for mobile devices based on an attention mechanism provided by the present invention improves the recognition accuracy of the model through data augmentation, significantly reduces the training cost, and moderately enhances the generalization ability. The use of a dynamic differential attention mechanism reduces the noise interference in the model, and by learning the identity-related information in the data, the scene coverage, accuracy, and efficiency of the model are improved. The use of model distillation combined with device performance, by extracting features from the teacher model and applying KMeans for clustering, greatly reduces the performance requirements for the device and improves the performance of the model used for authentication. Description of the Drawings
[0041] Figure 1 It is a flowchart of an implicit authentication method for mobile devices based on an attention mechanism of the present invention;
[0042] Figure 2 It is a graph showing the change of the TPR index in Experiment 1 of the present invention;
[0043] Figure 3 It is a graph showing the change of the TNR index in Experiment 1 of the present invention;
[0044] Figure 4 It is a graph showing the change of the ACC index in Experiment 1 of the present invention;
[0045] Figure 5This is the accuracy test result graph where the training set in Experiment 2 of the present invention consists of walking data and the test set consists of sitting data;
[0046] Figure 6 This is the accuracy test result graph where the training set in Experiment 2 of the present invention consists of sitting data and the test set consists of walking data;
[0047] Figure 7 This is the accuracy test result graph where the training set in Experiment 2 of the present invention consists of walking data and the test set also consists of walking data;
[0048] Figure 8 This is the accuracy test result graph where the training set in Experiment 2 of the present invention consists of sitting data and the test set also consists of sitting data;
[0049] Figure 9 This is the trend graph of the performance change of the index TPR in Experiment 3 of the present invention;
[0050] Figure 10 This is the trend graph of the performance change of the index TNR in Experiment 3 of the present invention;
[0051] Figure 11 This is the trend graph of the performance change of the index ACC in Experiment 3 of the present invention;
[0052] Figure 12 This is the graph of the change in memory usage of the student model and the teacher model in the entire 5 - second running time in Experiment 3 of the present invention. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0055] In response to the problems in the prior art, the present invention proposes an implicit authentication method for mobile devices based on the attention mechanism. After enhancing the data, a dynamic differential attention mechanism is used to improve the accuracy and coverage rate of the model, and at the same time, model distillation is performed so that the model can be deployed on more devices.
[0056] Such as Figure 1As shown, the implicit authentication method for mobile devices based on the attention mechanism in this embodiment includes the following steps:
[0057] Step 1: Collect multi-source data generated during the use of the mobile device based on multiple motion sensors of the mobile device and perform preprocessing.
[0058] Mobile devices are usually equipped with multiple motion sensors. Selecting appropriate sensor data is crucial for designing a system that can not only ensure a good user experience but also perform authentication. Considering the three aspects of device compatibility, user privacy and security, and environmental robustness, this embodiment selects the acceleration sensor, gravity sensor, and gyroscope to collect data. At the same time, considering the significant impact of power consumption on the user experience, two data collection modes are adopted: idle mode and active mode. To minimize power consumption, no sampling is performed in the idle mode. In the active mode, the sampling frequency is set to 50Hz, and a single data collection lasts for 3 seconds.
[0059] When the sensors collect data, various factors may introduce different types of noise. Differences in device usage and complex environments may lead to irregular noise, which will significantly affect the accuracy of model training. To ensure the quality of the data, preprocessing operations are performed on the data. The preprocessing operations in this embodiment first standardize the data and then remove the machine noise, human noise, etc. contained in the data.
[0060] Step 2: Use the generative-discriminative framework to perform data augmentation on the preprocessed multi-source data to obtain multi-source augmented data.
[0061] To effectively utilize the limited sensor data in mobile devices and improve the performance and robustness of the model, this embodiment adopts data augmentation. Mobile user authentication relies on the collection of sensor data, which may face the problem of insufficient data in specific scenarios, resulting in insufficient model training. For example, some behavioral patterns are more common, while rare patterns (such as abnormal operations or uncommon behaviors) have fewer data samples. The data augmentation method in this embodiment uses a conditional sampling strategy to generate synthetic samples for rare patterns, thereby supplementing the unbalanced data set. This helps the teacher model better handle diverse users and behaviors and improve its ability to identify minorities or abnormal operations. The data augmentation method in this embodiment is implemented based on the generative-discriminative framework, and this generative-discriminative framework (abbreviated as AUTHGANS) includes a generator, a transformer encoder, and a discriminator.
[0062] To address the complexity of mobile device authentication, in this embodiment, sensor data is collected based on a 3 - second time window and a 50Hz sampling frequency. Each time window contains 150 samples, and the data shape is (150, 3). In data augmentation and teacher model training, the generator uses 100 - dimensional Gaussian noise and 70 one - hot labels to generate realistic sensor data through a multi - layer perceptron (MLP). The data generated by the generator is processed by a transformer encoder. Specifically, it first undergoes upsampling and pixel flipping operations, and then is adjusted to the shape that conforms to the sensor format (150×3) as the input data for the discriminator. The discriminator maps the input data to a higher - dimensional space. After being processed by convolution and the transformer encoder and combined with the labels, it makes a final discrimination to determine the authenticity of the data. This generative - discriminative framework effectively compensates for the diversity of device conditions and user behaviors, improving the robustness and accuracy of the authentication method in different devices, scenarios, and user environments. Finally, the multi - source augmented data in this embodiment is uniformly represented as a two - dimensional vector [R, C], where R = 9 (the number of sensor axes) and C = 75 (the number of samples in the data segment), providing a convenient data format for subsequent model training.
[0063] Step 3: Process the multi - source augmented data using a sliding window, and adopt a cross - attention mechanism to convert the multi - source augmented data within one window into fused features, and use the labeled fused features to train the teacher model.
[0064] To obtain better model performance, the teacher model in this embodiment uses a Transformer model. To meet the requirements of the mobile device authentication scenario, the structure of the traditional Transformer model is improved by abandoning its fixed structure and introducing a more flexible, efficient, and adaptable mechanism.
[0065] The global attention calculation complexity of the traditional Transformer model is high. In this embodiment, a sparse attention mechanism is introduced to only focus on the historical data most relevant to the current time step to reduce the computational amount. The specific implementation uses sliding window attention, and each time step only focuses on the data of its previous time steps to form a local attention window.
[0066] Mobile devices are usually equipped with multiple sensors (such as accelerometers, gyroscopes, GPS). To make full use of this multi - source data, this embodiment designs a multi - channel embedding module to embed the data of different sensors into a unified feature space and fuse them through a cross - attention mechanism. The specific process is as follows:
[0067] The augmented data of each sensor channel is converted into a feature vector through an independent embedding layer.
[0068] The cross-attention mechanism is adopted to fuse multi-source feature vectors into a single feature vector. The cross-attention mechanism can exchange information between different modalities. For example, the acceleration data can focus on the rotation information in the gyroscope data, thereby enhancing the perception ability of complex behaviors (such as turning and jumping). This multi-modal fusion mechanism enables the model to comprehensively understand user behavior from multiple dimensions, improving the accuracy and robustness of authentication.
[0069] The softmax attention mechanism of the traditional Transformer model is insufficiently adaptable to the dynamic noise in mobile device sensor data. Therefore, in this embodiment, a dynamic differential attention mechanism is designed to adaptively adjust the differential parameters in the attention calculation according to the noise level of the input data. The specific implementation is as follows:
[0070] Introduce a dynamic differential parameter in the attention calculation , whose value is calculated in real time by a lightweight noise estimation module. The noise estimation module analyzes the short-term variance or frequency domain characteristics of the sensor data. For the sensor data (such as accelerometer or gyroscope data), the variance is calculated within a fixed time window (for example, the most recent time steps). The larger the variance, the higher the noise level. Therefore, should be increased, and vice versa should be decreased. The is dynamically adjusted to optimize the noise reduction effect. For example, in this embodiment, the multi-source enhanced data within a window is taken, and the variance of the enhanced data corresponding to each motion sensor is calculated. If the variance of the enhanced data of one motion sensor is greater than the reasonable fluctuation range, the dynamic differential parameter is increased; otherwise, the dynamic differential parameter is decreased. The dynamic differential parameter . The reasonable fluctuation range can be a preset fixed range or a dynamic range extracted from historical sensor enhanced data. And when increasing or decreasing the dynamic differential parameter, it can be increased or decreased according to a fixed gradient value. Therefore, the attention calculation formula in this embodiment is:
[0071]
[0072] In the formula, is the attention value, is the query vector, is the key vector, is the value vector, is the softmax function, is the transpose of the key vector , is the key vector dimensions. This adaptive mechanism enables the model to maintain high precision in different noise environments (such as user movement or noisy environments), significantly enhancing the robustness of authentication.
[0073] In addition, the rotation position encoding of the traditional Transformer model has a large computational overhead and is not suitable for mobile devices. Therefore, in this embodiment, this mechanism is abandoned, and the method of generating position embedding vectors based on rotation position encoding in the Transformer model is replaced with the method of generating position embedding vectors based on learned position encoding, capturing the position information in the time series through a small number of trainable parameters. The calculation formula for generating position embedding vectors based on learned position encoding proposed in this embodiment is:
[0074]
[0075] In the formula, is the position embedding vector, is the th time step, , is the total number of time steps, is the Embedding function.
[0076] Furthermore, the RMSNorm (Root Mean Square Normalization) mechanism in the traditional Transformer model uses fixed normalization parameters and is difficult to adapt to the changing input data distribution on mobile devices. Therefore, in this embodiment, an adaptive normalization mechanism is designed, enabling the model to dynamically adjust the feature distribution according to the characteristics of the input data by introducing trainable scaling and offset parameters. The specific formula is:
[0077]
[0078] In the formula, is the normalized eigenvalue output for the input feature , is the number of elements in the input feature , is the th element in the input feature , is a constant, is the scaling ratio, is the offset. and can adaptively scale and offset the eigenvalue, enhancing the generalization ability of the model under different users and environments, and can be pre-assigned according to empirical values or experiments.
[0079] The improved Transformer model in this embodiment shows significant advantages in the mobile device identity authentication scenario: Improved authentication accuracy: Compared with the traditional Transformer model, the new model's authentication accuracy in a noisy environment has increased by 20%. Especially when the user behavior is complex or the environment is changeable, it performs more stably. Optimized computational efficiency: Through sparse attention and learned positional encoding, the computational complexity of the model has been reduced by 30%, and the inference time has been shortened by 25%, ensuring fast response on mobile devices. Enhanced multi-modal perception: Multi-channel embedding and cross-attention mechanisms enable the model to fuse various sensor data, comprehensively understand user behavior, and reduce the false authentication rate by 15%.
[0080] Step 4: Obtain the student model by distilling the trained teacher model. When distilling, dynamically select the distillation position of the teacher model according to the performance score of the mobile device.
[0081] In lightweight mobile devices, computing resources such as processors, memory, and batteries are usually limited. And teacher models usually need to process complex multi-modal sensor data and involve large-scale feature processing. In such a resource-constrained environment, larger deep learning models may not be able to run efficiently, especially considering the relatively low processing and hardware capabilities of most lightweight devices on the current market. There are also differences in processing capabilities between devices, which means that deploying the same model on different hardware may face performance bottlenecks. In addition, the authentication process requires a quick response to ensure a smooth user experience, which is particularly important for applications with high real-time requirements. In this system, this embodiment proposes a model distillation strategy AuthFusion, and the specific steps are as follows:
[0082] Step 4.1: Calculate the performance score of the mobile device.
[0083] Traditional model distillation methods usually perform knowledge transfer at fixed layer positions. This approach ignores the differences in model layer requirements for different devices. Especially in the mobile device identity authentication scenario, the diversity of device performance requires the model to be able to flexibly adjust the knowledge transfer position. For this reason, this embodiment introduces a dynamic layer selection mechanism, which adaptively selects the distillation position in the teacher model according to the device capabilities, thereby optimizing the model's performance on different devices.
[0084] Before model deployment, the system will comprehensively evaluate the hardware capabilities of the device to determine the strength of its computing resources, and calculate a device capability score based on this, which serves as the basis for subsequent dynamic adjustment. The specific implementation is as follows:
[0085] High-performance devices (such as flagship smartphones): They receive higher scores due to their fast processors and sufficient memory. Low-performance devices (such as entry-level devices): They receive lower scores due to limited hardware resources. The following hardware parameters are mainly considered during scoring: Processor speed: Measures the computing power of a mobile device; Memory size: Determines the amount of data that a mobile device can process; Battery capacity: Affects the battery life and performance stability of a mobile device.
[0086] For example, obtain information such as processor speed, memory, and battery size according to the Android official API, and then calculate the performance score of the mobile device according to the formula: Performance score = 0.5 × Normalized processor speed value + 0.3 × Normalized memory size value + 0.2 × Normalized battery capacity value.
[0087] Step 4.2. Suppose the teacher model has layers. If the performance score is less than the scoring threshold, then take the first layers of the teacher model as the distillation range; otherwise, take the last layers of the teacher model as the distillation range.
[0088] Based on the performance score of the mobile device, dynamically determine which layer or layers of the teacher model to obtain knowledge hints from to find the best balance between performance and efficiency. Among them, high-performance devices select deeper hint positions in the teacher model. Deeper hints can provide richer and more complex feature information, which helps to improve the accuracy of the model in the identity authentication task. The powerful computing power of high-performance devices can be fully utilized to ensure the maximization of authentication accuracy; while low-performance devices select shallower hint positions in the teacher model. The features extracted by shallower hints have lower complexity and smaller computational burden, which can reduce resource consumption. It is suitable for resource-constrained devices to ensure smooth and efficient model operation.
[0089] Taking a 12-layer neural network as an example, if the performance score is less than the scoring threshold, then this mobile device is a low-performance device, and take the first 7 layers (layers 1 - 7) of the teacher model as the distillation range; otherwise, this mobile device is a high-performance device, and take the last 5 layers (layers 8 - 12) of the teacher model as the distillation range.
[0090] Step 4.3. For the distillation range, use the KMeans clustering method for clustering, and take the centroid of each cluster in the clustering result as the distillation position.
[0091] When the traditional KMeans clustering method selects the distillation position, it relies on the calculation of a fixed number of clusters and lacks direct optimization of the knowledge transfer effect. Therefore, this embodiment designs adaptive hint distillation. By introducing a feedback mechanism, the number of clusters is dynamically adjusted to maximize the performance of the student model. The specific implementation is as follows:
[0092] Based on the maximum and minimum values of the performance score, three score ranges are divided: a high score range, a medium score range, and a low score range.
[0093] According to the calculated performance score of the mobile device, determine the score range of the mobile device.
[0094] If the score range of the mobile device is the high score range, determine that the number of clusters is the first numerical value.
[0095] Alternatively, if the score range of the mobile device is the medium score range, determine that the number of clusters is the second numerical value.
[0096] Alternatively, if the score range of the mobile device is the low score range, determine that the number of clusters is the third numerical value; and the first numerical value is greater than the second numerical value, and the second numerical value is greater than the third numerical value.
[0097] For example, set the maximum value of the performance score to 1.0 and the minimum value to 0.0. Divide 0.8 - 1.0 (including 1.0) as the high score range, 0.5 - 0.8 (including 0.8) as the medium score range, and 0.0 - 0.5 (including 0.0 and 0.5) as the low score range. If the calculated performance score of the mobile device is in the high score range, take the number of clusters as 10. In the case of more clusters, the accuracy is higher; if the calculated performance score of the mobile device is in the medium score range, take the number of clusters as 5, which is a medium number of clusters, balancing accuracy and speed; if the calculated performance score of the mobile device is in the low score range, take the number of clusters as 3, which is fewer clusters, with speed being the priority.
[0098] During the training process of the student model, regularly evaluate its performance on the validation set (such as authentication accuracy, inference time) to determine whether the student model is trained to convergence.
[0099] In addition, to further reduce the computational complexity, in the layer representation of the teacher model, randomly sample a part of the data points for clustering instead of using all the data, that is, randomly sample the layers within the distillation range, and based on the random sampling results, use the KMeans clustering method for clustering.
[0100] At the same time, cache the clustering results on the mobile device to avoid repeated calculations and improve the response speed of the model.
[0101] Step 4.4: Train and output the student model based on the distillation position in the teacher model.
[0102] In this embodiment, by introducing mechanisms such as dynamic layer selection and adaptive prompt distillation, AuthFusion successfully optimizes the model distillation process, making it more suitable for the requirements of mobile device authentication scenarios. This method not only improves the efficiency and accuracy of knowledge transfer but also enhances the deployment adaptability of the model on resource-constrained devices, providing a more efficient, flexible, and robust solution for real-time authentication on mobile devices.
[0103] Step 5: Input the fusion features to be verified into the student model to obtain the authentication result output by the student model, where the authentication result is an illegal user or a legal user.
[0104] For the generation method of the fusion features to be verified, refer to the generation process of the fusion features in Steps 1 - 3, which will not be elaborated in this embodiment. Additionally, although this embodiment introduces Steps 1 - 5 in sequence, during the training process of the student model, Steps 1 - 4 can be executed independently, while during the application process of the student model, Step 5 can be executed independently.
[0105] This embodiment further conducts experiments to intuitively demonstrate the advantages of the solution in this embodiment.
[0106] Experiment 1: Energy efficiency of AUTHGANS.
[0107] To evaluate the impact of AUTHGANS in this embodiment on the training accuracy of the model, the experiment is carried out for training and testing on a dataset. The model is trained with denoising data using training sets with different numbers of users (the numbers of users are 100, 200, 300, 400, and 500 respectively), where the number of users refers to the number of users who collect sensor data in the dataset. For example, 100 means that the sensor data of 100 users is collected. As Figures 2 - 4 shown, two models are tested using the evaluation metrics of the test set with the same number of users (the number of users is 100). As the scale of the training set expands, the TPR (True Positive Rate), TNR (True Negative Rate), and ACC (Accuracy) of the model trained using AUTHGANS generally show an upward trend, while the TNR of the model without using AUTHGANS shows a decline. This indicates that as the data volume increases, the data augmentation effect of AUTHGANS becomes more significant, which can further improve the model accuracy. However, when the evaluation metrics approach the limit value of 100%, the marginal effect begins to appear.
[0108] From this experiment, it can be concluded that when the test set is smaller than the training set, the performance improvement brought by AUTHGANS increases with the expansion of the training set. It is particularly worth noting that AUTHGANS significantly improves the TNR index because after the model is retrained with data augmentation, its ability to identify unauthorized users is strengthened. Based on this, it can be judged that AUTHGANS data augmentation can effectively utilize limited data to improve the model efficiency.
[0109] Experiment 2: Energy efficiency of the dynamic differential attention mechanism.
[0110] To evaluate the improvement of the model coverage rate after adding the dynamic differential attention mechanism in different scenarios, the experiment selected a variety of data sets and used the dynamic differential attention mechanism and the traditional attention mechanism for comparative analysis of the model coverage rate. In these data sets, the experiment trained the model using "Walk" data and "Sit" data respectively, and used data in another state for testing. Based on the source of the data, the coverage ability of the model in different states was evaluated. The experimental objects were the dynamic differential attention mechanism (Diff Attention), the traditional attention mechanism (Attention), and the LSTM neural network (LSTM) proposed in the present invention. The experiment selected four different brand mobile phones, namely A, B, C, and D, for sensor data collection.
[0111] As Figures 5 - 8 shown, although the accuracy of the model decreases when using data in different states for testing, the model adopting the dynamic differential attention mechanism can still achieve an accuracy rate of more than 90%, which is sufficient to meet the user's needs.
[0112] From this experiment, it can be concluded that by introducing the dynamic differential attention mechanism, even if only data of a single user behavior is used for training, the model can still maintain a high accuracy rate. This significantly improves the coverage ability of the model in different scenarios, thus ensuring a better user experience.
[0113] Experiment 3: Energy efficiency of the model distillation strategy AuthFusion.
[0114] In this experiment, the model trained based on the generative-discriminative framework in Experiment 1 was used as the teacher model, and the AuthFusion distillation method was applied to obtain the student model. First, the sizes of the student model and the teacher model were compared. Table 1 shows the sizes of the teacher models trained under training sets with different numbers of users (the numbers of users are 100, 200, 300, 400, and 500 respectively) and the sizes of their distilled student models. The results show that AuthFusion significantly reduces the scale of the model.
[0115] Table 1 The sizes of the teacher models trained with different numbers of users and the sizes of their distilled student models
[0116]
[0117] Next, the performance of the models before and after distillation was evaluated. The teacher model and the student model were trained using a training set with the same number of users (100 users), and the two models were tested using the evaluation metrics of test sets with different numbers of users (100, 200, 300, 400, and 500 users). As Figures 9 - 11 shown, although the accuracy of the model after distillation decreased slightly, it still remained within an acceptable range, ensuring high accuracy while safeguarding user privacy and security.
[0118] In addition, the student model and the teacher model trained with 100 users were selected for deployment testing, and the memory usage during the authentication process was recorded. As Figure 12 shown, the distilled student model had obvious advantages over the teacher model in terms of memory usage and recognition time.
[0119] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0120] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. An implicit authentication method for mobile devices based on attention mechanism, characterized in that: The attention mechanism-based implicit authentication method for mobile devices includes: Collect and pre-process multi-source data generated when the mobile device is used based on multiple motion sensors of the mobile device; The preprocessed multi-source data is enhanced using the generation-discrimination framework to obtain multi-source enhanced data; Use sliding windows to process multi-source enhanced data, and use a cross-attention mechanism to convert the multi-source enhanced data in a window into fused features, and use the labeled fused features to train the teacher model; The student model is obtained by distilling the trained teacher model. During distillation, the distillation position of the teacher model is dynamically selected according to the performance score of the mobile device. Inputting the fusion feature to be verified into the student model, obtaining the identity authentication result output by the student model, wherein the identity authentication result is an illegal user or a legal user; The teacher model is a Transformer model, which adopts a dynamic differential attention mechanism. The attention calculation formula of the dynamic differential attention mechanism is: ; In the formula, is the attention value, is the query vector, is the key vector, is a value vector, is the softmax function, is the key vector The transpose of is the key vector The dimension of is a dynamic differential parameter, the dynamic differential parameter The adjustments are as follows: Take the multi-source enhanced data within a window and calculate the variance of the enhanced data corresponding to each motion sensor; If the variance of the enhanced data of a motion sensor is larger than the reasonable fluctuation range, increase the dynamic difference parameter ; otherwise reduce the dynamic difference parameter , the dynamic differential parameter ; The step of obtaining a student model by distilling the trained teacher model includes: Calculate performance scores for mobile devices; Let the teacher model have a total of If the performance score is less than the score threshold, the teacher model's previous The layer is used as the distillation scope; otherwise, the posterior layer of the teacher model is taken Layers serve as distillation ranges; The distillation range is clustered using the KMeans clustering method, and the centroid of each cluster in the clustering result is used as the distillation position; Train and output a student model based on the distilled positions in the teacher model.
2. According to claim 1, the implicit identity authentication method for mobile devices based on the attention mechanism is characterized in that: The proposed generation-discrimination framework includes a generator, a transformer encoder, and a discriminator.
3. The method for implicit identity authentication of mobile devices based on attention mechanism according to claim 1, characterized in that: The Transformer model's method of generating a position embedding vector based on rotational position coding is replaced by a method of generating a position embedding vector based on learning position coding. The calculation formula for generating a position embedding vector based on learning position coding is: ; In the formula, is the position embedding vector, For the time steps, , is the total time step, It is the Embedding function.
4. The method for implicit identity authentication of mobile devices based on the attention mechanism according to claim 1, characterized in that: The calculation formula of the RMSNorm normalization mechanism of the Transformer model is as follows: ; In the formula, For input features The normalized eigenvalues of the output, For input features The number of elements in , is the input feature The elements, is a constant, is the scaling factor, is the offset.
5. The method for implicit identity authentication of mobile devices based on attention mechanism according to claim 1, characterized in that: The number of clusters for the KMeans clustering method is determined as follows: Based on the maximum and minimum values of the performance score, three score ranges are obtained: high score range, medium score range, and low score range; Determining a score range of the mobile device according to the calculated performance score of the mobile device; If the score range of the mobile device is a high score range, determining the number of clusters to be a first number value; Alternatively, if the score range of the mobile device is a medium score range, determining the number of clusters to be a second number value; Alternatively, if the score range of the mobile device is a low score range, the number of clusters is determined to be a third number value; and the first number value is greater than the second number value, and the second number value is greater than the third number value.
6. The method for implicit identity authentication of mobile devices based on the attention mechanism according to claim 1, characterized in that: The distillation range is clustered using a KMeans clustering method, including: randomly sampling layers within the distillation range; Based on the random sampling results, KMeans clustering method is used for clustering.
Citation Information
Patent Citations
Multi-modal fusion wavelet knowledge distillation video behavior identification method and system based on cross attention
CN115294498A
Bimodal biometric feature recognition network model training method
CN116385832A