Implicit identity authentication method based on multi-motion sensor data and related device
By constructing a multi-channel Transformer network to process data from multiple motion sensors, the problem of low authentication accuracy in existing technologies is solved, achieving more efficient implicit identity authentication and improving user experience and security.
Patent Information
- Application Number
- CN202511091720.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-11
AI Technical Summary
Existing implicit identity authentication methods based on deep learning struggle to accurately capture the multidimensional semantic and temporal dynamic characteristics of multi-motion sensor data, resulting in low authentication accuracy.
A multi-channel Transformer network is used to process motion sensor data. Temporal features and channel features are extracted by constructing a first-channel and a second-channel Transformer network, respectively, and then combined using a weighted fusion method and input into a classifier for identity authentication.
It improves the accuracy of implicit identity authentication, enhances authentication capabilities in different scenarios, and improves user experience and security.
Smart Images

Figure CN120934815A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of identity authentication technology, specifically relating to an implicit identity authentication method and related apparatus based on multi-motion sensor data. Background Technology
[0002] With the continuous advancement of communication and mobile internet technologies, smartphones have become deeply integrated into people's daily lives, becoming an indispensable tool in modern society. Smartphones have not only dramatically transformed how information is accessed and how social interaction occurs (e.g., news browsing and instant messaging), but also provided people with diverse functions such as convenient lifestyle services (e.g., online shopping, navigation, and food delivery), financial services (e.g., mobile payments and stock investment), and entertainment and media consumption (e.g., online games and video viewing). While the widespread use of smartphones has made users' lives more efficient, the large amount of personal privacy data and sensitive information stored on smart devices (e.g., chat logs, contacts, text messages, emails, bank account information, and passwords) is at risk of malicious access. Once attackers gain unauthorized access, it could lead to serious consequences such as privacy breaches, financial losses, criminal activity, and identity theft.
[0003] To effectively protect personal privacy data on smart devices from malicious attacks and intrusions, smart devices currently employ various authentication mechanisms to ensure the legitimacy of the individual's identity when using them. Early authentication methods were primarily knowledge-based, typically requiring users to provide private information upon initial setup, which was then verified each time they logged in. These methods assume that only legitimate users share knowledge with the authentication system, thus ensuring that others cannot easily access the smart device without authorization. For example, traditional knowledge-based authentication methods include passwords and pattern locks. While these methods have low implementation costs and high convenience, allowing for rapid deployment in most systems, they still have some significant limitations. First, knowledge-based authentication carries the risk of password leakage. Traditional password authentication systems are vulnerable to brute-force attacks, spying attacks, or social engineering attacks. For example, attackers can use automated tools to repeatedly try passwords from common password databases, thereby brute-forcing passwords and gaining unauthorized access. Second, these authentication methods often provide a poor user experience. Knowledge-based authentication methods require users to remember multiple passwords or change them regularly to ensure account security. This increases the user's memory burden and severely impacts the user experience. For example, in password authentication scenarios, users often need to frequently enter complex and hard-to-remember passwords, especially when the screen of a smart device is small, making the input experience even more cumbersome.
[0004] To overcome the aforementioned limitations, biometric authentication technologies have developed rapidly. This technology typically involves a smart device collecting and storing samples of the user's physiological information (such as fingerprints and facial images) during the user registration phase, which are then used for subsequent authentication. During the authentication phase, the smart device collects the current user's physiological information (also called physiological features) and compares it with the features stored during registration to confirm the user's identity. Currently widely used physiological features include facial recognition, fingerprints, palm prints, and iris scans.
[0005] While biometric authentication methods have addressed some of the issues of brute-force attacks and poor user experience in traditional cryptographic authentication methods, they still face certain challenges and limitations: First, biometric authentication is vulnerable to forgery. While everyone's physiological characteristics are unique, some physiological information is easily copied or forged in real life. Facial recognition technology may face "face spoofing attacks," where attackers can use facial photos found on social media to impersonate victims and bypass the authentication system of smart devices. Second, biometric authentication technology is costly. To collect users' physiological information, smart devices may need specialized hardware, such as fingerprint scanners, iris scanners, or high-performance facial recognition cameras. The procurement and installation costs of these devices are high, potentially hindering widespread adoption by ordinary users. Furthermore, biometric authentication methods are typically only applicable to specific entry points, such as identity verification for unlocking and payment, and cannot provide continuous authentication throughout the entire use of a smart device. This means that once a user completes authentication, they gain long-term access to the smart device, increasing security risks. If the device is lost or stolen, attackers could use the already authenticated information to gain full access to the smart device within a short period.
[0006] Behavioral biometrics-based implicit authentication offers an innovative and promising approach to protecting the security of smart devices. Unlike traditional authentication methods (such as passwords or fingerprint recognition), behavioral biometrics authenticates users based on their unique behavioral patterns (such as touchscreen usage, input methods, and movement trajectories) during interaction with the smart device. The biggest advantage of this method lies in its implicit authentication feature; the authentication process is automatic and seamless, requiring no active user intervention. This provides a convenient and effective means of protecting smart device security, effectively preventing unauthorized access.
[0007] Among numerous implicit identity authentication solutions based on behavioral biometrics, motion dynamics biometric methods based on motion sensor data have attracted widespread attention. These methods utilize motion sensors such as accelerometers and gyroscopes built into smart devices to capture the dynamic motion characteristics of users during interactions with the device. Research shows that individuals exhibit unique motion dynamic characteristics when performing different activities (such as walking and cycling), which can therefore be used for identity authentication. Because this sensor data is continuous and dynamic, accurate identity authentication can be performed by analyzing motion patterns without requiring any additional user intervention, significantly improving user experience and enhancing security.
[0008] Current methods primarily focus on using deep learning techniques to model motion sensor data, automatically learning the complex relationship between user behavior and identity through data-driven approaches. Feature extraction is considered a critical step affecting authentication accuracy. Traditional methods typically employ manual feature extraction techniques to extract statistical features (such as variance, mean, and entropy) from motion sensor data to characterize dynamic changes in motion patterns. These methods rely on domain expert knowledge for feature selection and extraction, resulting in high computational costs and difficulty in capturing all potential and complex dynamic patterns from large datasets.
[0009] However, with the successful application of deep learning technology in fields such as computer vision and natural language processing, researchers have begun to use deep learning methods to replace traditional manual feature extraction. These methods learn behavioral patterns from sensor data in an automated manner, saving significant time spent on manual exploration and improving the efficiency and accuracy of feature extraction. Nevertheless, existing deep learning-based authentication methods still face two main challenges: 1) It is difficult to capture information about the sequential changes of multiple channels from multiple motion sensors. Some researchers have proposed treating multi-motion sensor data as images and using two-dimensional convolutional neural networks (CNNs) to extract motion dynamics features from multiple channels. However, this approach, which processes multi-channel data into images, ignores the multi-dimensional semantics of the multi-motion sensor data itself. Unlike the semantic consistency of two-dimensional images, the temporal and channel dimensions of multi-motion sensor data have different semantics. Therefore, simply using CNNs to convolve multi-dimensional data may not accurately capture motion dynamics features, thus affecting the performance of authentication systems.
[0010] 2) Ignore the time dynamics of each channel dimension To address the problems identified in Challenge 1), some researchers have begun to treat multi-motion sensor data as time series, employing sequential models (such as RNNs and LSTMs; RNN stands for Recurrent Neural Network, and LSTM stands for Long Short-Term Memory) to capture features exhibiting changes over time. However, these methods neglect the interrelationships between cross-channel temporal features (i.e., time-channel features), resulting in an inability to effectively capture changes in motion patterns across different operational scenarios (e.g., standing, sitting, and lying down), thus leading to low accuracy in implicit identity authentication. The usage patterns and resulting motion signals of smart devices vary significantly across different scenarios; changes in cross-channel temporal features can reveal specific operational scenarios. Summary of the Invention
[0011] The purpose of this invention is to provide an implicit identity authentication method and related device based on multi-motion sensor data, in order to solve the problem of low accuracy of implicit identity authentication in the prior art.
[0012] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides an implicit identity authentication method based on multi-motion sensor data, comprising the following steps: Acquire data from several types of motion sensors; The obtained motion sensor data of several types are standardized to obtain the standardized motion sensor data of several types. The standardized motion sensor data of several types is segmented to obtain several fixed-length data segments, and these fixed-length data segments are further divided into time-series data and channel data. Construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder. The time series data and channel data are respectively input into the first channel Transformer network and the second channel Transformer network to obtain the time series features and channel features; The temporal features and channel features are fused to obtain the fused features; The fused features are input into the classifier to obtain the implicit identity authentication result.
[0013] A further improvement of the present invention is that the various types of motion sensor data include accelerometer sensor data, gyroscope sensor data, magnetometer sensor data, and rotation vectormeter sensor data.
[0014] A further improvement of this invention is that the standardization process is a normalization process, and the calculation formula for the normalization process is:
[0015] in, and motion sensors The mean and standard deviation of the channel. For the original motion sensor Measurements in the time series of the channel, motion sensor after normalization Measurements in the time series of the channel.
[0016] A further improvement of the present invention is that the segmentation process is specifically as follows: a sliding window mechanism is used to segment the standardized motion sensor data of several types to obtain several data segments of fixed length.
[0017] A further improvement of the present invention is that the fusion is specifically: a weighted fusion method is used to fuse the temporal features and the channel features to obtain the fused features.
[0018] A further improvement of this invention is that the calculation formula for weighted fusion is as follows: ( ) in, The final fused features, As a time series feature, As a channel feature, and These are the attention weights for temporal features and channel features, respectively. This is a vector concatenation operation; Attention weights for temporal features Attention weights for channel features The calculation formula is: , = ( ) in, This is a standardized exponential function used to transform a vector into a probability distribution. Features after initial fusion; Features after initial fusion The calculation formula is:
[0019] in, The weight matrix is a learnable linear transformation. As a time series feature, As a channel feature, This is a learnable bias term.
[0020] Secondly, the present invention provides an implicit identity authentication system based on multi-motion sensor data, including a data acquisition module, a standardization processing module, a segmentation processing module, a network construction module, a feature extraction module, a feature fusion module, and an implicit identity authentication module; The data acquisition module is used to acquire several types of motion sensor data; The standardization processing module is used to standardize the obtained motion sensor data of several types to obtain standardized motion sensor data of several types. The segmentation processing module is used to segment several types of motion sensor data after standardization to obtain several fixed-length data segments, and then divide the fixed-length data segments into time-series data and channel data. The network construction module is used to construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder. The feature extraction module is used to input time series data and channel data into the first channel Transformer network and the second channel Transformer network respectively to obtain time series features and channel features; The feature fusion module is used to fuse temporal features and channel features to obtain fused features; The implicit identity authentication module is used to input the fused features into the classifier to obtain the implicit identity authentication result.
[0021] Thirdly, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the implicit authentication method based on multi-motion sensor data described above.
[0022] Fourthly, the present invention provides a storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the implicit authentication method based on multi-motion sensor data described above.
[0023] Compared with the prior art, the present invention has the following beneficial effects: The implicit authentication method based on multi-motion sensor data proposed in this invention has two main advantages. First, it acquires several types of motion sensor data, resulting in diverse data and improving the accuracy of subsequent implicit authentication. Second, it constructs a first-channel Transformer network and a second-channel Transformer network, inputting temporal data and channel data into these networks respectively to obtain temporal and channel features. This operation not only improves the targeting of feature extraction but also captures the characteristics of motion sensor data from different perspectives, avoiding the low accuracy of implicit authentication caused by extracting a single feature. Therefore, it effectively solves the problem of low accuracy in implicit authentication in existing technologies. Attached Figure Description
[0024] Figure 1 This is a flowchart of the implicit identity authentication method based on multi-motion sensor data of the present invention; Figure 2 This is a schematic diagram of the implicit identity authentication system based on multi-motion sensor data of the present invention; Figure 3 This is an architecture diagram of the implicit identity authentication method based on multi-motion sensor data in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the gated dual-tower Transformer fusion network in Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the structure of the multi-channel one-dimensional CNN network in Embodiment 3 of the present invention; Figure 6 This is a schematic diagram of the structure of the electronic device of the present invention. Detailed Implementation
[0025] To further understand the content of this invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are merely illustrative and not limiting of the invention.
[0026] Example 1: The flowchart of the implicit authentication method based on multi-motion sensor data of this invention is as follows: Figure 1 As shown, the implicit identity authentication method based on multi-motion sensor data of the present invention includes the following steps: S1. Acquire data from several types of motion sensors; S2. Standardize the obtained motion sensor data of several types to obtain standardized motion sensor data of several types; S3. The standardized motion sensor data of several types is segmented to obtain several fixed-length data segments, and these fixed-length data segments are divided into time-series data and channel data; S4. Construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder. S5. Input the time series data and channel data into the first channel Transformer network and the second channel Transformer network respectively to obtain the time series features and channel features; S6. Fuse the temporal features and channel features to obtain the fused features; S7. Input the fused features into the classifier to obtain the implicit identity authentication result.
[0027] Example 2: A schematic diagram of the implicit identity authentication system based on multi-motion sensor data of the present invention is shown below. Figure 2 As shown, the implicit identity authentication system based on multi-motion sensor data of the present invention includes a data acquisition module, a standardization processing module, a segmentation processing module, a network construction module, a feature extraction module, a feature fusion module, and an implicit identity authentication module.
[0028] The data acquisition module is used to acquire data from several types of motion sensors.
[0029] The standardization processing module is used to standardize the obtained motion sensor data of several types, resulting in standardized motion sensor data of several types.
[0030] The segmentation processing module is used to segment several types of motion sensor data after standardization to obtain several fixed-length data segments, and then divides these fixed-length data segments into time-series data and channel data.
[0031] The network construction module is used to construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder.
[0032] The feature extraction module is used to input time series data and channel data into the first channel Transformer network and the second channel Transformer network, respectively, to obtain time series features and channel features.
[0033] The feature fusion module is used to fuse temporal features and channel features to obtain fused features.
[0034] The implicit identity authentication module is used to input the fused features into the classifier to obtain the implicit identity authentication result.
[0035] Example 3: The architecture diagram of the implicit authentication method based on multi-motion sensor data of this invention is as follows: Figure 3 As shown, the implicit identity authentication method based on multi-motion sensor data of the present invention includes the following steps: S1. Acquire data from several types of motion sensors.
[0036] First, several types of motion sensor data are acquired through a smart terminal device (the acquisition cycle is once every 2 seconds). These types of motion sensor data include accelerometer sensor data, gyroscope sensor data, magnetometer sensor data, and rotation vector meter sensor data.
[0037] The accelerometer sensor data, gyroscope sensor data, magnetometer sensor data, and rotation vector meter sensor data are all triaxial measurements, specifically represented as follows: , , , , , , , , , , and .
[0038] in, This represents the triaxial measurements from the motion sensor. , , and These represent the accelerometer, gyroscope, magnetometer, and rotary vectormeter, respectively.
[0039] S2. Standardize the obtained motion sensor data of several types to obtain standardized motion sensor data of several types.
[0040] In this step, the standardization process is a normalization process, and the calculation formula for the normalization process is:
[0041] in, and motion sensors The mean and standard deviation of the channel (also called the axis). For the original motion sensors (accelerometers, gyroscopes, magnetometers, and rotation vector meters). Measurements in the time series of the channel, motion sensor after normalization Measurements in the time series of the channel.
[0042] motion sensor Channels and motion sensors The formula for channel normalization is referenced from that of motion sensors. aisle.
[0043] S3. The standardized motion sensor data of several types is segmented to obtain several fixed-length data segments, and these fixed-length data segments are further divided into time-series data and channel data.
[0044] This step specifically employs a sliding window mechanism to segment the standardized motion sensor data of several types, resulting in several fixed-length data segments. The process is described in detail below: To segment the standardized motion sensor data of several different types, a fixed window length is used on each motion sensor channel. (Window length) A sliding window (with a value range of 0.5-5 seconds) is used to generate a series of overlapping data segments. .
[0045] in, This indicates the number of data segments, with * indicating a specific channel of each motion sensor, and a sliding step size of ΔS. When the value of ΔS is greater than or equal to the fixed window length... When the conditions are met, the generated data segments will not overlap; otherwise, the data segments will overlap.
[0046] S4. Construct the first-channel Transformer network and the second-channel Transformer network.
[0047] S5. Input the time series data and channel data into the first channel Transformer network and the second channel Transformer network respectively to obtain the time series features and channel features.
[0048] S6. Fuse the temporal features and channel features to obtain the fused features.
[0049] In this embodiment, steps S4, S5, and S6 are treated as a whole to construct a gated dual-tower Transformer fusion network. A schematic diagram of the gated dual-tower Transformer fusion network is shown below. Figure 4 As shown.
[0050] In this embodiment, the gated dual-tower Transformer fusion network is trained using a fully supervised learning approach. In fully supervised learning, the gated dual-tower Transformer fusion network uses a labeled dataset to guide its training process. The label information provides the network with the desired output, enabling the gated dual-tower Transformer fusion network to progressively adjust its parameters to minimize the difference between the predicted results and the true labels. The specific training process of the gated dual-tower Transformer fusion network is described below: (1) Training sample preparation Given a set of user training samples, each sample contains a corresponding label (i.e., the correct prediction result). These labels provide the target output for the gated dual-tower Transformer fusion network, guiding it to optimize its prediction capabilities during training.
[0051] (2) Backpropagation After each forward propagation, the difference between the predictions of the gated dual-tower Transformer fusion network and the true labels is calculated. The gradient of the loss function with respect to the network parameters is then calculated using the backpropagation algorithm. By calculating the gradient, the backpropagation algorithm identifies the error contribution of each parameter in the gated dual-tower Transformer fusion network, thereby helping the network update its parameters.
[0052] (3) Parameter update Based on the calculated gradient, the parameters of the gated dual-tower Transformer fusion network are updated in the opposite direction of the gradient. This process is achieved using gradient descent, allowing the predictions of the gated dual-tower Transformer fusion network to gradually approach the true labels. To accelerate the training process and improve the convergence of the gated dual-tower Transformer fusion network, the Adam optimizer is used. The Adam optimizer combines the advantages of momentum and adaptive learning rates, effectively handling sparse gradients and large-scale data.
[0053] (4) Loss function During the training of the gated dual-tower Transformer fusion network, the binary cross-entropy loss function is used as the objective function. The binary cross-entropy loss function is commonly used in classification problems, especially in binary classification tasks, to guide the optimization process by quantifying the difference between the model output and the true label. In this embodiment, this loss function helps the gated dual-tower Transformer fusion network learn how to distinguish between legitimate users and imposters, improving the accuracy of subsequent implicit authentication.
[0054] Step S4 will be explained in detail below: First, construct the first-channel Transformer network and the second-channel Transformer network.
[0055] The first channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder, while the second channel Transformer network uses a multi-channel one-dimensional convolutional neural network (also called a multi-channel one-dimensional CNN) and a Transformer encoder (also called a Transformer).
[0056] A schematic diagram of the structure of a multi-channel one-dimensional convolutional neural network is shown below. Figure 5 As shown, a multi-channel one-dimensional convolutional neural network consists of a series of one-dimensional filters and pooling layers. When performing convolution operations on one-dimensional temporal data, the filters extract features from the sequence through a window that moves along the time axis. The size of the convolution kernel defines the number of time steps involved in each convolution operation. Similar to two-dimensional CNNs, 1D CNNs (also called multi-channel one-dimensional convolutional neural networks) can not only use pooling layers to reduce the dimensionality of feature maps, but also design various 1D CNN architectures by using combinations of basic convolution and pooling operations. The input of a one-dimensional CNN model is a data segment, and the output is a representation of that data segment.
[0057] In a multi-channel one-dimensional convolutional neural network, each convolutional layer has 128 convolutional kernels, which are used to extract features from the input one-dimensional data. Each convolutional kernel performs a sliding window operation along the time axis or sequence direction on the one-dimensional input data to extract features at a specific location. After each convolution operation, a pooling operation is performed to capture features at different scales, reducing output complexity and preventing overfitting. For each feature vector, max pooling is used to select the most significant feature value. By alternately applying the above convolution and pooling operations, the one-dimensional CNN model can transform the input data segment into a representation vector.
[0058] The Transformer encoder (or Transformer for short) consists of two sub-layers: a multi-head self-attention network and a feedforward network. Furthermore, each sub-layer is followed by residual connections and normalization operations to generate the output representation. The details of these two sub-layers are described below: A. Multi-head self-attention network In Transformer, a multi-head attention mechanism is designed to capture the complex relationships between data segments. Self-attention is the core component of the multi-head self-attention module, used to capture data segments. It demonstrates the dependencies and long-distance associations between different positions. This allows the model to model the interactions and relationships between elements at different positions within a single sequence, without relying on traditional recursive mechanisms. This mechanism allows each element to compute its own representation based on the representations of all other elements.
[0059] The self-attention calculation process is explained below: a. Use the representation of each element (obtained through linear transformation) as input for Query, Key, and Value, respectively; b. Calculate the attention weights by computing the dot product of the query and all keys; c. Perform scaling and Softmax operations to obtain the attention score of each element to all other elements.
[0060] For each sub-data segment representation It and The attention weights between the representations of all other sub-data segments are calculated using a scaled dot product attention mechanism, and the calculation formula is as follows:
[0061] in, This indicates that it is obtained by performing a linear transformation on the input matrix. This represents the attention weights between segments of the set. (Used...) The scaling factor is used to avoid the inner product result from becoming too large.
[0062] d. Multiply each value by its corresponding attention weight, sum all the weighted values, and obtain the representation of the current element.
[0063] Multi-head self-attention mechanisms (also called multi-head self-attention networks) have the following advantages: First, multi-head self-attention allows the model to learn different relationships in the sequence from different perspectives, providing richer and more diverse representation capabilities. Second, different heads focusing on different parts of the sequence can more accurately capture various features and patterns in the sequence. Moreover, the multi-head mechanism enhances the learning ability for different types of information, which helps improve the model's generalization ability.
[0064] Assuming there is a bullish self-attention Size. The self-attention segment vectors will be concatenated to form the output data segment representation of the Transformer from the input data segment. Specifically, given the representation matrix of the Transformer (L-1) layer... and The nth parallel self-attention module, the nth The calculation method for each attention point is as follows:
[0065] in, , The linear transformation weights for this layer.
[0066] The final formula for calculating the multi-head self-attention output is:
[0067] in, To output the projection matrix, This is a vector concatenation operation.
[0068] B. Feedforward Neural Network For each sub-layer in the Transformer, the output of multi-head self-attention is fed into the feedforward neural network to enhance the interaction between different dimensions. This network consists of two transformation functions: a linear transformation and a ReLU non-linear transformation. These linear and ReLU non-linear transformations transform the output of the previous layer... Convert to the current layer's representation The calculation formula is:
[0069] in, (Feedforward Neural Network) The weight matrix is the first-level linear transformation. This is the weight matrix for the second-level linear transformation. and This is the bias vector.
[0070] In step S6, the temporal features and channel features are fused to obtain the fused features (also called gating feature fusion).
[0071] This step specifically employs a weighted fusion method to fuse temporal features and channel features, resulting in fused features.
[0072] The formula for weighted fusion is: ( ) in, The final fused features, As a time series feature, As a channel feature, and These are the attention weights for temporal features and channel features, respectively. This is a vector concatenation operation.
[0073] Attention weights for temporal features Attention weights for channel features The calculation formula is: , = ( ) in, This is a standardized exponential function used to transform a vector into a probability distribution. These are the features after the initial fusion. Will and Stacked along the third dimension; the Softmax operation is performed along the third dimension, ensuring... + =1, achieving normalized weighting.
[0074] Features after initial fusion The calculation formula is:
[0075] in, The weight matrix is a learnable linear transformation. As a time series feature, As a channel feature, This is a learnable bias term.
[0076] S7. Input the fused features into the classifier to obtain the implicit identity authentication result.
[0077] In this step, the classifier is an OC-SVM classifier. The output of the OC-SVM classifier is a confidence score ∈ [0,1], with the following meaning: When the confidence score approaches 1, it indicates that the behavior is highly consistent with normal user behavior and the authentication status is normal.
[0078] When the confidence score is lower than the set threshold, it indicates that the user's behavior may be abnormal and further verification is required.
[0079] In this step, the kernel function parameter γ of the OC-SVM classifier is set to 0.1, and the parameter ν is set to 0.05. The settings of γ and ν are used to control the generalization ability of the OC-SVM classifier so as to better authenticate implicit identities.
[0080] To verify the effectiveness of the implicit authentication method based on multi-motion sensor data proposed in this invention, this embodiment constructs five datasets containing different combinations of motion sensors based on the collected data, as follows: Ac only contains accelerometer data; AcGy includes accelerometer and gyroscope data; AcGyMa includes data from accelerometers, gyroscopes, and magnetometers; AcGyOr contains data from accelerometers, gyroscopes, and rotation vectormeters; AcGyMaOr contains data from accelerometers, gyroscopes, magnetometers, and rotation vectormeters.
[0081] Based on five datasets containing different combinations of motion sensors, this embodiment compares the method of the present invention with six existing mainstream identity authentication methods. The comparison results are shown in Table 1.
[0082] Table 1 Comparison Results
[0083] As shown in Table 1, integrating data from four types of sensors (accelerometer, gyroscope, magnetometer, and attitude orientation) yields the best authentication performance. The FAR (False Acceptance Rate), FRR (False Rejection Rate), and EER (Equal Error Rate) of the method in this invention are 8.3%, 9.4%, and 8.8%, respectively. When using only accelerometer and gyroscope data, the performance (FAR, FRR, and EER) drops to 10.3%, 11.1%, and 10.7%, respectively. This data indicates that different sensors are complementary in behavior modeling, and their combined use can significantly improve the accuracy of implicit identity authentication.
[0084] Example 4: Please see Figure 6 As shown, the present invention also provides an electronic device 100 based on an implicit authentication method using multiple motion sensor data; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.
[0085] The memory 101 can be used to store the computer program 103. The processor 102 implements the steps of the implicit authentication method based on multi-motion sensor data described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device 100 (such as audio data), etc. In addition, the memory 101 may include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other non-volatile solid-state storage device.
[0086] The at least one processor 102 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or any conventional processor. The processor 102 is the control center of the electronic device 100, connecting various parts of the electronic device 100 via various interfaces and lines.
[0087] The memory 101 in the electronic device 100 stores multiple instructions to implement an implicit authentication method based on multi-motion sensor data, and the processor 102 can execute the multiple instructions to achieve the following: Acquire data from several types of motion sensors; The obtained motion sensor data of several types are standardized to obtain the standardized motion sensor data of several types. The standardized motion sensor data of several types is segmented to obtain several fixed-length data segments, and these fixed-length data segments are further divided into time-series data and channel data. Construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder. The time series data and channel data are respectively input into the first channel Transformer network and the second channel Transformer network to obtain the time series features and channel features; The temporal features and channel features are fused to obtain the fused features; The fused features are input into the classifier to obtain the implicit identity authentication result.
[0088] Example 5: If the modules / units integrated in the electronic device 100 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, and a read-only memory (ROM).
[0089] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0090] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0091] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. An implicit authentication method based on multi-motion sensor data, characterized in that, Includes the following steps: Acquire data from several types of motion sensors; The obtained motion sensor data of several types are standardized to obtain the standardized motion sensor data of several types. The standardized motion sensor data of several types is segmented to obtain several fixed-length data segments, and these fixed-length data segments are further divided into time-series data and channel data. Construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network uses a multi-channel one-dimensional convolutional neural network and a Transformer encoder. The time series data and channel data are respectively input into the first channel Transformer network and the second channel Transformer network to obtain the time series features and channel features; The temporal features and channel features are fused to obtain the fused features; The fused features are input into the classifier to obtain the implicit identity authentication result.
2. The implicit authentication method based on multi-motion sensor data according to claim 1, characterized in that, The various types of motion sensor data include accelerometer sensor data, gyroscope sensor data, magnetometer sensor data, and rotation vectormeter sensor data.
3. The implicit authentication method based on multi-motion sensor data according to claim 1, characterized in that, The standardization process is a normalization process, and the calculation formula for the normalization process is: in, and motion sensors The mean and standard deviation of the channel. For the original motion sensor Measurements in the time series of the channel, motion sensor after normalization Measurements in the time series of the channel.
4. The implicit authentication method based on multi-motion sensor data according to claim 1, characterized in that, The segmentation process specifically involves using a sliding window mechanism to segment the standardized motion sensor data of several types into several fixed-length data segments.
5. The implicit authentication method based on multi-motion sensor data according to claim 1, characterized in that, The fusion specifically involves using a weighted fusion method to fuse temporal features and channel features to obtain fused features.
6. The implicit authentication method based on multi-motion sensor data according to claim 5, characterized in that, The formula for weighted fusion is: ( ) in, The final fused features, As a time series feature, As a channel feature, and These are the attention weights for temporal features and channel features, respectively. This is a vector concatenation operation; Attention weights for temporal features Attention weights for channel features The calculation formula is: , = ( ) in, This is a standardized exponential function used to transform a vector into a probability distribution. Features after initial fusion; Features after initial fusion The calculation formula is: in, The weight matrix is a learnable linear transformation. As a time series feature, As a channel feature, This is a learnable bias term.
7. The implicit authentication method based on multi-motion sensor data according to claim 1, characterized in that, The classifier is an OC-SVM classifier.
8. An implicit identity authentication system based on multi-motion sensor data, characterized in that, It includes a data acquisition module, a standardization processing module, a segmentation processing module, a network construction module, a feature extraction module, a feature fusion module, and an implicit identity authentication module; The data acquisition module is used to acquire several types of motion sensor data; The standardization processing module is used to standardize the obtained motion sensor data of several types to obtain standardized motion sensor data of several types. The segmentation processing module is used to segment several types of motion sensor data after standardization to obtain several fixed-length data segments, and then divide the fixed-length data segments into time-series data and channel data. The network construction module is used to construct a first-channel Transformer network and a second-channel Transformer network; the first-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder, and the second-channel Transformer network adopts a multi-channel one-dimensional convolutional neural network and a Transformer encoder. The feature extraction module is used to input time series data and channel data into the first channel Transformer network and the second channel Transformer network respectively to obtain time series features and channel features; The feature fusion module is used to fuse temporal features and channel features to obtain fused features; The implicit identity authentication module is used to input the fused features into the classifier to obtain the implicit identity authentication result.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the implicit authentication method based on multi-motion sensor data as described in any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the implicit authentication method based on multi-motion sensor data as described in any one of claims 1 to 7.