An implicit identity authentication method based on multimodal feature fusion

Through the implicit identity authentication method of multimodal feature fusion, the problem of insufficient flexibility and adaptability of single modal technology in the face of user behavior and biometric changes is solved, and higher recognition accuracy and system stability are achieved.

CN119653361BActive Publication Date: 2025-05-16UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510185176.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-05-16
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

The single-modal implicit identity authentication technology lacks flexibility and adaptability in the face of changes in user behavior and biometrics, resulting in a decrease in identification accuracy and system stability and security.

Method used

The implicit identity authentication method based on multimodal feature fusion is adopted to extract the features of sensor data and touch screen data through convolutional neural networks and GRU deep learning networks, and deep feature fusion is used to use feature channel exchange module, adaptive weight feature fusion module and cross-modal Mamba module for deep feature fusion, and finally use binary support vector machines for identity authentication.

Benefits of technology

It improves the adaptability and robustness of the model, enhances the accuracy and reliability of the recognition, and can effectively identify biometric features and identify identity in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119653361B_ABST
    Figure CN119653361B_ABST
Patent Text Reader

Abstract

The present invention provides an implicit identity authentication method based on multimodal feature fusion, which converts the discrete signal of the acquired sensor data into a grayscale image, then uses a pre-trained convolutional neural network to extract features, combines GRU to extract touch screen sequence data features, and then sends the features of the two modalities into a multimodal fusion network for feature fusion, including preliminary feature fusion and deep feature fusion. The preliminary feature fusion performs preliminary fusion of the features of the two modalities through an adaptive feature fusion module and a channel exchange feature fusion module, respectively, and the deep feature fusion performs multimodal deep feature fusion through an improved network model CrossMamba, and generates the final fused features, and finally performs identity authentication through an identity authentication module. The flexibility and adaptability of the model are enhanced through the scheme of the present invention, and the accuracy and reliability of identity authentication are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of implicit identity authentication, and in particular to an implicit identity authentication method based on multimodal feature fusion. Background Art

[0002] According to statistics released by Gartner, a world-renowned research organization, 1.348 billion smartphones were sold worldwide in 2020. Due to the rapid development of smartphones and related industries, today's society has entered the mobile Internet era from the PC (Personal Computer) era. With the continuous popularization of smartphones and the continuous enrichment of their functions, users are accustomed to storing personal privacy such as pictures, voice, chat records and financial information in their mobile phones, making mobile phones an important target for malware and hacker attacks. Therefore, it is necessary for academia and industry to develop a reliable user authentication mechanism on smartphones. Token-based authentication is generally based on some physical media (such as ID cards) to provide users with reliable authentication, but physical devices are inconvenient to carry and users find it difficult to log in after the device is lost. Therefore, token-based authentication is rarely used in smartphones.

[0003] Knowledge-based authentication does not require users to carry physical devices with them. It is a widely used authentication method in many daily scenarios and one of the most commonly used authentication methods for smartphones. However, knowledge-based authentication technology has certain limitations, which can be attributed to external and internal factors. External factors mainly come from attackers and malware technology, such as attackers can obtain passwords and PINs through direct observation, or guess passwords by analyzing screen stains to access user resources; internal factors mainly come from user usage habits, such as users find it difficult to remember complex random passwords, and may use a combination of numbers such as phone numbers or birthdays as passwords, or choose extremely simple passwords such as "123456". Token-based and knowledge-based authentication have the above limitations, while biometric authentication technology does not require users to carry physical devices and does not require users to remember extra information. Therefore, biometric authentication technology has become a research hotspot in the field of mobile phone user authentication.

[0004] Biometric authentication technology can be divided into two categories: the first category is authentication technology based on human physiological characteristics, and the second category is authentication technology based on human behavioral characteristics. Authentication technology based on physiological characteristics uses human physiological characteristics to authenticate the user's identity. Common technologies include fingerprint, face, iris and retina recognition. Currently, fingerprint recognition and face recognition are the most widely used in commercial applications.

[0005] The advantages of fingerprint recognition systems are that they are easy to implement, have high recognition accuracy and are fast. However, fingerprint patterns are susceptible to cutting, dirt and wear and can be easily copied. Attackers can log in by constructing fingerprint films through grease analysis. Face recognition systems train models by learning user facial features and use the learned models to authenticate user identities. The security risk of face recognition is that attackers can crack the authentication system by constructing GIF images of legitimate user facial photos or using 3D masks. In addition, both face and fingerprint recognition can only achieve one-time authentication at the login entrance, and their static authentication characteristics also make them easy to copy.

[0006] Authentication based on behavioral characteristics realizes identity authentication by detecting the user's behavior. There are two main forms: the first is to identify the user's identity by collecting and analyzing the behavioral habits reflected by the specific operations of the user when using the mobile phone, such as analyzing the applications frequently used by the user and the cellular base stations frequently connected; the second is to identify the user's identity based on the user's actual hand and body movements, such as gait, screen touch and keystroke when walking. Authentication technology based on behavioral characteristics does not require users to physically carry and remember additional information, and it is difficult for others to know, record and imitate, thus avoiding the defects of the above-mentioned authentication technology based on physiological characteristics. Compared with the fingerprint and face recognition methods that perform one-time authentication at the login entrance, the authentication method based on behavioral characteristics can achieve continuous and transparent and non-intrusive identity authentication to the user, and has therefore become a research hotspot in the field of smartphone user identity authentication.

[0007] Single-modal implicit authentication technologies usually rely on a specific type of user behavior characteristics (such as typing patterns, mouse movement trajectories) or biometric characteristics (such as fingerprints, facial features) to authenticate identity, and the expressiveness of these characteristics is precisely limited by their own variability and stability. For example, typing patterns may change due to user fatigue, mood changes, or device changes (such as keyboard type); similarly, facial recognition may also be affected by the user's expression, makeup, lighting conditions, or age. When these characteristics change, the single-modal authentication system itself is often unable to adapt to the new feature changes in a timely manner, resulting in a significant decrease in its recognition accuracy.

[0008] Specifically, a single-modality authentication system relies on learning and matching a fixed pattern of a certain feature. If the user's feature deviates from the fixed pattern expected by the system, the system will not be able to accurately identify the user, which may result in a high false positive rate (identifying a legitimate user as an unauthorized user) or a missed positive rate (misidentifying an unauthorized user as a legitimate user). This reliance on a single feature results in the system's lack of flexibility and adaptability when faced with changes in user behavior and biometrics, which directly affects its overall authentication effectiveness and reliability.

[0009] Traditional machine learning methods require a separate model for training because the system needs to process and analyze a large amount of data from different modalities. This may cause excessive waiting time during the user registration phase, resulting in longer training time; the reasoning time becomes longer because during the authentication process, the data of each modality needs to be processed and analyzed by its corresponding model, and then the output results are integrated. This multi-step process leads to longer reasoning time, and users may need to wait a long time to complete the authentication, which directly affects the user experience. In addition, traditional machine learning models are usually trained on fixed data sets and lack adaptability to new or unseen feature changes, which leads to a decrease in model recognition accuracy and affects the stability and security of the system.

[0010] At the same time, because traditional machine learning methods are based on manual features and feature selection relies on expert knowledge, the extracted features may be subjectively biased and difficult to cope with unseen data changes, thus affecting the generalization ability of the model; and manual features are difficult to capture the complexity of multimodal data, which makes the model unable to effectively process data from different sources, further reducing the authentication performance. In addition, manual feature extraction cannot fully utilize the advantages of big data, because this method often relies heavily on the feature set preset by experts, resulting in a limited feature space and unable to capture all useful information in the data, especially in high-dimensional and multimodal data, important features may be omitted, which limits the performance improvement of the model on large-scale data sets. Because these methods have poor adaptability and are difficult to cope with dynamic changes, the model needs to be updated frequently, increasing maintenance costs. Finally, these processes usually involve complex calculations, resulting in low computational efficiency, especially in application scenarios that require fast response. Summary of the invention

[0011] In order to solve the above problems, the present invention proposes an implicit identity authentication method based on multimodal feature fusion, which solves the problems existing in single-modal data authentication technology, and uses the behavioral feature data when the user interacts with the mobile phone to perform additional identity authentication on the current user after the user unlocks the mobile phone using the main authentication method. The scheme of the present invention can fully capture the user's behavior pattern, and the model has stronger adaptability, robustness and accuracy, and can effectively identify biometric features and perform identity discrimination in different scenarios.

[0012] The present invention provides an implicit identity authentication method based on multimodal feature fusion, comprising:

[0013] Step S1: preprocessing the acquired sensor data and touch screen data, and converting the preprocessed sensor data into a grayscale image;

[0014] Step S2: Use the designed convolutional neural network to extract features from the grayscale image obtained by converting sensor data, and use the GRU deep learning network to extract features from touch screen data;

[0015] Step S3: Use the feature channel exchange module and the adaptive weight feature fusion module to perform preliminary feature fusion on sensor data and touch screen data;

[0016] Step S4: Shallow feature convolution fusion, that is, perform channel number fusion and dimensionality reduction on the input features;

[0017] Step S5: Use the cross-modal Mamba module for deep feature fusion;

[0018] Step S6: Use the features after deep fusion for user identity authentication to verify its legitimacy.

[0019] Furthermore, the step S1 includes:

[0020] Step S11: First, perform zero-value filling operations on the acquired sensor data and touch screen data, then perform wavelet denoising on the filled data, and then perform Kalman filtering;

[0021] Step S12: Normalize the sensor data processed in step S11, and then calculate the GASF matrices of the sensor on the three coordinate axes x, y, and z respectively;

[0022] Step S13: Linearly transform the GASF matrix values to the interval [0,1], and then adjust each element value in the matrix to the range of 0 - 255;

[0023] Step S14: Generate a pseudo deBrujin sequence;

[0024] Step S15: Let the time series length be N, including K time steps, and then according to the generated deBrujin sequence, select P consecutive time steps of GASF matrices for matrix splicing to obtain the final GASF matrix, where P < K and P can be divisible by K. Finally, convert the spliced GASF matrix into a grayscale image.

[0025] Furthermore, the convolutional neural network in the step S2 includes 9 convolutional layers and 3 max pooling layers. The 9 convolutional layers are divided into three convolutional modules, namely convolutional module one, convolutional module two, and convolutional module three. Each convolutional module includes three convolutional layers. The entire convolutional neural network is divided into three convolutional modules and a fully connected layer part;

[0026] Convolutional module one: For the input single-channel image , where, represents a real number set, H and W represent the height and width of the image respectively. The three layers of convolution module 1 extract the 32-dimensional features of the input image in turn, and the size of each convolution kernel is , with a step size of 1 and a padding of 1, expressed as:

[0027] ;

[0028] in, Indicates Features extracted by layer convolution, and Respectively The weights and biases of the layer convolution, * indicates a two-dimensional convolution operation, is the ReLU activation function, and a 2×2 maximum pooling operation is performed after the third convolution layer to halve the spatial resolution:

[0029] ;

[0030] Represents the maximum pooling operation, and the output feature map size becomes ;

[0031] Convolutional module 2: Input , the three-layer convolution of convolution module 2 expands the number of feature channels to 64, and the convolution operation is expressed as:

[0032]

[0033] in, and Respectively The weights and biases of the layer convolution, , after the sixth convolution layer, a 2×2 maximum pooling is performed to further reduce the resolution:

[0034]

[0035] The output feature map size becomes ;

[0036] Convolutional module 3: Input , the three-layer convolution of convolution module 3 expands the number of feature channels to 128, and the convolution operation is expressed as:

[0037]

[0038] in, and Respectively The weights and biases of the layer convolution, , after the ninth convolution layer, a 2×2 maximum pooling is performed:

[0039]

[0040] The output feature map size is ;

[0041] Fully connected layer: After convolution and pooling operations, the feature map is flattened into a vector:

[0042]

[0043] The output feature map size is , and input the obtained output into the fully connected layer:

[0044]

[0045]

[0046] in, , , represents the weight of the fully connected layer, the output dimension is 256, and the final output is a 256-dimensional feature vector: .

[0047] Furthermore, the step S3 comprises:

[0048] Step S31: Input the sensor feature vector and the touch screen feature vector extracted in step S2 into the feature channel exchange module to obtain a preliminarily fused feature vector , specifically:

[0049] For the input sensor feature vector and the touch screen feature vector ,in , is the dimension of the vector, and Divided into two sub-vectors:

[0050]

[0051]

[0052] in, It's the first half. and It is the second half, and then the truncated features are spliced:

[0053]

[0054] Among them, Concat represents the vector concatenation operation, , , and finally obtained ;

[0055] Step S32: The sensor feature vector extracted in step S2 and the touch screen feature vector Input the adaptive weight feature fusion module to obtain the initially fused feature vector , expressed as:

[0056]

[0057] in, and is the fusion weight, satisfying the constraints: .

[0058] Furthermore, step S4 includes: performing shallow feature convolution fusion on the two initially fused features obtained in step S3, that is, for the input and ,renew and :

[0059]

[0060] ;

[0061] Will and Spliced ​​into a dimension The feature of B represents batch_size, which is the amount of data processed by the model at one time during training. The concatenated features are convolved to reduce the input dimension from 512 to 256, and then the original Add it back to the convolution result, and then update the Again and Perform splicing, perform convolution extraction after splicing, and update , and the residual connection is also used to Add the convolution result.

[0062] Furthermore, the cross-modal Mamba module in step S5, namely the CrossMamba module, includes a normalization layer Norm, a multi-layer perceptron MLP, a one-dimensional convolution module C1, an activation function Activation, a classical space state model SSM and a function for reconstructing the shape of the feature matrix , the workflow is:

[0063] First, the channel exchange feature And weight fusion features Normalize: , ;

[0064] The normalized sequence is then input into the MLP, using the linear layer Perform a linear mapping to map the sequence to a higher dimension P:

[0065]

[0066]

[0067]

[0068] , Further features are extracted through one-dimensional convolution and activation function:

[0069]

[0070]

[0071] Then based on , Generate space state parameters:

[0072]

[0073]

[0074] in, , is the state dynamic matrix, , is the input control matrix, , To output the projection matrix, use SSM , Perform dynamic modeling and get output , :

[0075]

[0076]

[0077] Using gating factors Perform weighted activation:

[0078]

[0079]

[0080] The activated modal features are fused and the residual term is added:

[0081]

[0082] Get the final fusion features.

[0083] Furthermore, the step S6 includes: using a binary support vector machine to perform identity verification. After obtaining the output features of step S5, the support vector machine is initialized using a Gaussian kernel function, and the Gaussian kernel function is:

[0084] ,

[0085] represents the sample points in the input space, It is a hyperparameter. Through the Gaussian kernel function binary support vector machine, the optimal hyperplane can be found in the high-dimensional space, so as to effectively judge the identity. If the identity is legal, the device can continue to be used, otherwise the device is prohibited from being used.

[0086] The beneficial technical effects of the present invention are as follows: first, in terms of feature extraction, the modal conversion of multimodal data is used to perform deep feature extraction, that is, the discrete signal of the acquired sensor data is converted into a grayscale image, which contains a specific numerical sequence and can better represent the characteristics of the data signal involved. Then, by using transfer learning, a specific pre-trained convolutional neural network model is used for feature extraction to obtain higher-dimensional and more expressive features, which fully utilizes the advantages of big data and greatly improves the model performance and adaptability; secondly, for the collected multimodal data, a network model based on a convolutional neural network is designed to perform multimodal feature fusion, including shallow feature fusion and deep feature fusion. The shallow feature fusion performs preliminary fusion of modal data through adaptive feature fusion and channel exchange, and enhances the preliminary interaction and expression ability of cross-modal data; the deep feature fusion uses a cross-modal Mamba module, namely the CrossMamba module, for deep fusion, and intends to fully fuse the multimodal data features and generate comprehensive features after fusion, which are rich in layers and more expressive, and make the authentication system more stable, flexible and adaptable, greatly enhancing the accuracy and reliability of recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0087] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0088] Figure 1It is a flow chart of an implicit identity authentication method based on multimodal feature fusion provided by an embodiment of the present invention;

[0089] Figure 2 It is a channel exchange feature fusion process provided by an embodiment of the present invention;

[0090] Figure 3 is the adaptive weight feature fusion process provided by the embodiment of the present invention, wherein and represents the fusion weight;

[0091] Figure 4 is a structural diagram of CorssMamba provided by an embodiment of the present invention, wherein: represents the channel switching characteristics, represents the weight fusion feature, is the final fusion feature. DETAILED DESCRIPTION

[0092] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0093] The present invention proposes an implicit identity authentication method based on multimodal feature fusion, such as Figure 1 As shown, the method comprises the following steps:

[0094] Step S1: preprocessing the acquired sensor data and touch screen data, and converting the preprocessed sensor data into a grayscale image;

[0095] Step S2: Use the designed convolutional neural network to extract features from the grayscale image converted from the sensor data, and use the GRU deep learning network to extract features from the touch screen data;

[0096] Step S3: using a feature channel exchange module and an adaptive weight feature fusion module to perform preliminary feature fusion on the sensor data and the touch screen data;

[0097] Step S4: shallow feature convolution fusion, that is, the input features are fused and the number of channels is reduced;

[0098] Step S5: Use the cross-modal Mamba module to perform deep feature fusion;

[0099] Step S6: Use the deeply fused features to authenticate the user and verify its legitimacy.

[0100] In the whole process, the time series data is first converted into image processing in order to use CNN to extract deeper and richer expressive features contained in the time series data, and on this basis, the shallow and deep features extracted are fused;

[0101] In the entire network, the channel exchange feature fusion module and the adaptive weight feature fusion module initially fuse the features, preliminarily enhancing the interactivity of features of different modalities, and then further fuse them through shallow feature convolution to prepare for the subsequent deep feature fusion module; in the deep feature fusion module, multimodal features will fully interact between modal features in high-dimensional space to build a more comprehensive and robust feature representation;

[0102] The implementation of the present invention needs to be based on the mobile phone sensor data of all users and the touch screen data of all users.

[0103] The step S1 further comprises:

[0104] Step S11: First, the acquired sensor data and touch screen data are filled with zero values, and then the filled data are subjected to wavelet denoising, because the wavelet denoising operation can effectively process static noise, especially excels in noise removal in different frequency bands, and is suitable for processing the multi-scale structure of signals; then the Kalman filter is processed, which is good at processing dynamic systems, and can effectively track and predict signals in the time series changes of signals, and is suitable for dealing with situations with time-varying noise.

[0105] The signal noise of the accelerometer and gyroscope can be regarded as high-frequency random jitter, which conforms to the Gaussian white noise assumption. At the same time, the Kalman filter can effectively smooth the signal by dynamically adjusting the weights of the predicted value and the observed value. Combined with wavelet denoising, it can retain key dynamic change features (such as sudden acceleration or rotation), which is better than a simple low-pass filter. Therefore, it is suitable to use Kalman filtering and wavelet denoising to denoise the data of these two sensors. By combining the above operations, wavelet denoising can first suppress a large amount of noise in the frequency domain, and the Kalman filter can further dynamically track and adjust the denoised signal in the time series to ensure the accuracy of the signal changing over time.

[0106] Step S12: For the sensor data processed in step S11, normalize it first, and then calculate the GASF (Gramian Angular Summation Field) matrix of the sensor (accelerometer, gyroscope) on the three coordinate axes x, y, and z respectively. The GASF matrix is ​​a technology for converting time series into images. It is often used for visualization of time series data and input feature construction of deep learning models (such as convolutional neural networks). The core idea of ​​GASF is to map the time series to a polar coordinate system through angle transformation, and then use the Gram matrix to generate a two-dimensional matrix representation to retain the temporal relationship and global information of the time series.

[0107] Generation of GASF:

[0108] For sensor data Normalization is performed, where N represents the length of the time series. The value is scaled to Within the range:

[0109]

[0110] In the above formula, UB and LB represent the upper and lower limits of the sensor sequence data, that is, the maximum and minimum values ​​in the sequence; then each sequence value is encoded as a cosine, the timestamp is encoded as a radius, and the angle value corresponding to each timestamp point is calculated. , to generate the GASF matrix:

[0111]

[0112]

[0113] Among them, n represents the timestamp of each data, and L represents the constant factor of the Euclidean regularized polar coordinate system. By mapping the data into polar coordinates, the capture of the relative angle information between points can be enhanced, and the geometric meaning of each value in the time series can be intuitively represented.

[0114] Step S13: After obtaining the GASF matrix of the sequence data on multiple channels of the sensor, the matrix values ​​are linearly transformed to the interval [0,1], and then each element value in the matrix is ​​adjusted to the range of 0-255 (the pixel value of the image is an 8-bit unsigned integer (uint8), and its value range is [0,255]), that is, for each element of the matrix , first perform a linear transformation, then calculate :

[0115]

[0116]

[0117] Step S14: Generate a pseudo de Bruijn sequence. First, assign the numbers 0 to 5 to different signals in the generated sequence. For a de Bruijn sequence of order n (in this embodiment, n is 3, representing the tuple length) based on an alphabet of size k (in this embodiment, k is 6, representing the number of signals), it is a cyclic sequence in which each possible string of length n (i.e., a triple) based on the alphabet {0, 1, 2, 3, 4, 5} appears exactly once as a consecutive subsequence. Different from the traditional de Bruijn sequence, in the present invention, there is no need to include the permutations and combinations of triples that have already appeared in the sequence, nor triples containing repeated symbols, such as (3, 3, 3).

[0118] To construct the pseudo de Bruijn sequence, in this embodiment, a simple greedy algorithm is used to ensure that each newly added triple has not appeared before by minimizing the addition of signals. Although for a de Bruijn sequence with n = 3 and k = 6, the minimum length is 218, the length of the sequence generated by this method is 25. The finally obtained sequence is: 0, 1, 2, 3, 4, 5, 0, 2, 4, 5, 1, 3, 0, 4, 1, 2, 5, 3, 0, 2, 0, 5, 1, 3, 4. This sequence is regarded as a pseudo de Bruijn sequence.

[0119] Step S15: Assume that the length of the time series is N, including K time steps, and each time step contains M consecutive sequence data. For the x, y, and z axes of each sensor, taking a single time step as a unit, generate a GASF matrix of the size of a unit time step on each coordinate axis. That is, within the same time step, 6 GASF matrices of the same size will be obtained. Then, according to the generated de Bruijn sequence, select P consecutive time steps of GASF matrices for matrix splicing to obtain the final GASF matrix, where P < K and P is divisible by K. Then, convert the spliced GASF matrix into a grayscale image.

[0120] The step S2 further includes:

[0121] Step S21: The present invention uses a convolutional neural network structure including 9 convolutional layers and 3 maximum pooling layers to extract features from the grayscale image converted from the sensor data to obtain a 256-dimensional sensor feature vector. The convolutional neural network is used to extract feature vectors from the input single-pass value image. The network extracts features layer by layer, and the number of channels is gradually expanded from 32 to 128. At the same time, the spatial resolution is gradually reduced through the pooling operation. The whole process is expressed as:

[0122]

[0123] in, is the input single-channel image of size , respectively represent the height and width of the image; It is the convolution and pooling part, which is used to extract multi-level features. It is a fully connected part, which is used to map the convolution features to a vector space of fixed dimension. The 9 convolutional layers are divided into three convolutional modules, namely convolutional module 1, convolutional module 2 and convolutional module 3. Each convolutional module includes three convolutional layers. Therefore, the entire convolutional neural network can be divided into three convolutional modules and a fully connected layer:

[0124] Convolution module 1: For the input single-channel image ,in, represents a real number set, H and W represent the height and width of the image respectively, and three layers of convolution extract 32-dimensional features in turn. The size of each convolution kernel is , the step size is 1, the padding is 1, and the feature map size is guaranteed to remain unchanged, expressed as:

[0125]

[0126] in, Indicates Features extracted by layer convolution, and Respectively The weights and biases of the layer convolution, * indicates a two-dimensional convolution operation, is the ReLU activation function, and a 2×2 maximum pooling operation is performed after the third convolution layer to halve the spatial resolution:

[0127]

[0128] Represents the maximum pooling operation, and the output feature map size becomes .

[0129] Convolutional module 2: Input , the three-layer convolution expands the number of feature channels to 64, and the convolution operation is expressed as:

[0130]

[0131] in, and Respectively The weights and biases of the layer convolution, , after the 6th convolution layer, a 2×2 maximum pooling is performed to further reduce the resolution:

[0132]

[0133] The output feature map size becomes .

[0134] Convolutional module 3: Input , the three-layer convolution expands the number of feature channels to 128, and the convolution operation is expressed as:

[0135]

[0136] in, and Respectively The weights and biases of the layer convolution, , after the 9th convolution layer, a 2×2 maximum pooling is performed:

[0137]

[0138] The output feature map size is .

[0139] Fully connected layer: After convolution and pooling operations, the feature map is flattened into a vector:

[0140]

[0141] The output feature map size is , and input the obtained output into the fully connected layer:

[0142]

[0143]

[0144] in, , , represents the weight of the fully connected layer, the output dimension is 256, and the final output is a 256-dimensional feature vector:

[0145]

[0146] Step S22: Use GRU to extract features from the standardized touch screen data. GRU (Gated Recurrent Unit) is an improved variant of recurrent neural network (RNN) specially designed for processing sequence data. It is controlled by a gating mechanism, which can alleviate the problem of gradient disappearance in traditional RNN and show good results in many tasks. At the same time, the calculation efficiency is higher than LSTM (Long Short-Term Memory Network). By extracting features through the GRU network, a 256-dimensional touch screen feature vector is also obtained for subsequent multi-modal feature fusion.

[0147] The step S3 further comprises:

[0148] Step S31: Input the sensor feature vector and the touch screen feature vector extracted in step S2 into the feature channel exchange module to obtain a preliminarily fused feature vector , the fusion process is as follows Figure 2 As shown, specifically:

[0149] For the input sensor feature vector ,in is the dimension of the vector, and the touch screen feature vector ,in is the dimension of the vector, and Divided into two sub-vectors:

[0150]

[0151]

[0152] in, It's the first half. and is the second half. Then perform feature concatenation on the truncated features:

[0153]

[0154] Among them, Concat represents the vector concatenation operation, , , and finally obtained In the above steps, we get and The dimensions are all 256 (i.e. = ), so we finally get The dimension of is also 256.

[0155] Step S32: The sensor feature vector extracted in step S2 and the touch screen feature vector Input the adaptive weight feature fusion module to obtain the initially fused feature vector , the fusion process is as follows Figure 3 As shown, specifically:

[0156]

[0157] in, and is the fusion weight, satisfying the constraints:

[0158]

[0159] During the network training and convergence process, and It is a dynamically adjustable parameter that can be automatically adjusted through optimization objectives and gradually reach a balanced state.

[0160] generally, and Some form of weight update strategy (e.g., gradient descent-based learning) can be used for learning. This embodiment uses a learnable parameter To express The initial value of is calculated as follows and :

[0161]

[0162] This method ensures , and the weights are differentiable during training, so they can be optimized by back propagation. Finally, the initial fusion feature vector It will be used as input for subsequent modules.

[0163] Described step S4 specifically comprises:

[0164] Before performing deep feature fusion, the two initially fused features obtained in step S3 are fused by shallow feature convolution. and ,renew and :

[0165]

[0166]

[0167] Will and Spliced ​​into a dimension The feature of B represents batch_size, which is the amount of data processed by the model at one time during training. The concatenated features are convolved to reduce the input dimension from 512 to 256, and then the original Add it back to the convolution result, and then update the Again and Perform splicing, perform convolution extraction after splicing, and update , and the residual connection is also used to Add the convolution result.

[0168] The step S5 specifically includes:

[0169] Using the CrossMamba module, the features are deeply fused and the channel exchange features obtained in step S4 are And weight fusion features Input the designed CrossMamba module and obtain the fused features after full fusion. ,This module is intended to enhance the cross-modal expressiveness of features;

[0170] The CrossMamba module is a module for multimodal data fusion processing based on the initial Mamba model. It is a cross-modal sequence modeling module built based on the Mamba model, which is used for interactive feature fusion and modeling between two modalities. The structure of the CrossMamba module is as follows: Figure 4 As shown in the figure, Norm is the normalization layer, MLP module stands for multi-layer perceptron, which is responsible for the preliminary processing of the standardized feature values, C1 module represents the one-dimensional convolution operation on the input features, Activation is the activation function, and SSM is the classic spatial state model. is a function that reshapes the feature matrix.

[0171] In the calculation process of the CrossMamba module, the weight fusion feature As the main modal feature, firstly, the channel exchange feature And weight fusion features Normalize (Norm):

[0172] ,

[0173] The normalized sequence is then input into the MLP, using the linear layer Perform a linear mapping to map the sequence to a higher dimension P:

[0174]

[0175]

[0176]

[0177] , Further features are extracted through one-dimensional convolution and activation function:

[0178]

[0179]

[0180] Then based on , Generate space state parameters:

[0181]

[0182]

[0183] in, , is the state dynamic matrix, , is the input control matrix, , To output the projection matrix, use SSM , Perform dynamic modeling and get output , :

[0184]

[0185]

[0186] Using gating factors Perform weighted activation:

[0187]

[0188]

[0189] The activated modal features are fused and the residual term is added:

[0190]

[0191] So far, the final fusion features are obtained.

[0192] The step S6 further comprises:

[0193] Step S61: Perform user identity authentication to verify its legitimacy, extract the data features of the user to be authenticated, and input them into the authentication module for judgment. In the present invention, a binary support vector machine is used to perform identity legitimacy verification:

[0194] Support vector machine is a commonly used supervised learning method. Binary support vector machine is used to solve classification problems with only two categories (positive and negative) by finding an optimal hyperplane:

[0195]

[0196] in, is the weight vector, b is the bias term, x is the input sample point, and data point The classification is determined by the following rules:

[0197] , classified as positive class.

[0198] , classified as negative class.

[0199] Divide the data into two categories and maximize the classification margin:

[0200]

[0201] At the same time, the following constraints are met:

[0202]

[0203] in Indicates the category of the data point.

[0204] Specifically, after the output features of step S5 are obtained, the support vector machine is initialized using a Gaussian kernel function, which is:

[0205]

[0206] represents the sample points in the input space, It is a hyperparameter, and the appropriate γ is automatically selected according to the data size. Through the Gaussian kernel, the binary support vector machine can find the optimal hyperplane in the high-dimensional space, thereby effectively performing identity judgment. If the identity is legal, the device can continue to be used, otherwise the device is prohibited from being used.

[0207] The present invention proposes an implicit identity authentication method based on multimodal feature fusion, which solves the problems of poor adaptability, low recognition accuracy, insufficient feature extraction, etc. of the existing implicit recognition under a single modality. By integrating grayscale image conversion of time series data, self-designed feature extraction based on convolutional neural network, and implicit identity authentication of multimodal fusion authentication, the complexity is reduced while the authentication accuracy is greatly improved, thereby achieving fast and efficient identity recognition and authentication.

[0208] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An implicit identity authentication method based on multimodal feature fusion, characterized in that: The method includes: Step S1: Preprocess the acquired sensor data and touch screen data, and convert the preprocessed sensor data into a grayscale image; Step S2: Use the designed convolutional neural network to extract features from the grayscale image obtained by converting the sensor data, and use the GRU deep learning network to extract features from the touch screen data. The convolutional neural network includes 9 convolutional layers and 3 max pooling layers. The 9 convolutional layers are divided into three convolutional modules, namely convolutional module one, convolutional module two, and convolutional module three. Each convolutional module includes three convolutional layers. The entire convolutional neural network is divided into three convolutional modules and a fully connected layer part; Step S3: Use the feature channel exchange module and the adaptive weight feature fusion module to perform preliminary feature fusion on the sensor data and the touch screen data. Specifically: Step S31: Input the sensor feature vector and the touch screen feature vector extracted in step S2 into the feature channel exchange module to obtain a preliminarily fused feature vector , for the input sensor feature vector and the touch screen feature vector ,in , is the dimension of the vector, and Divided into two sub-vectors: ; ; in, It's the first half. and It is the second half, and then the truncated features are spliced: , where Concat represents the vector concatenation operation, , , and finally obtained ; Step S32: The sensor feature vector extracted in step S2 and the touch screen feature vector Input the adaptive weight feature fusion module to obtain the initially fused feature vector , expressed as: ,in, and is the fusion weight, satisfying the constraints: ; Step S4: shallow feature convolution fusion, that is, the input features are fused and dimensionally reduced in number of channels. Specifically, the two initially fused features obtained in step S3 are fused by shallow feature convolution. and ,renew and : ; ; Will and Spliced ​​into a dimension The feature of B represents batch_size, which is the amount of data processed by the model at one time during training. The concatenated features are convolved to reduce the input dimension from 512 to 256, and then the original Add it back to the convolution result, and then update the Again and Perform splicing, perform convolution extraction after splicing, and update , and the residual connection is also used to Add the convolution result; Step S5: Use the cross-modal Mamba module to perform deep feature fusion. The cross-modal Mamba module includes a normalization layer Norm, a multi-layer perceptron MLP, a one-dimensional convolution module C1, an activation function Activation, a classic space state model SSM, and a function for reconstructing the shape of the feature matrix. ; Step S6: Use the deeply fused features for user identity authentication to verify its legality.

2. The method according to claim 1, characterized in that The step S1 further includes: Step S11: First, perform zero-value filling operations on the acquired sensor data and touch screen data, then perform wavelet denoising on the filled data, and then perform Kalman filtering; Step S12: Normalize the sensor data processed in step S11, and then calculate the GASF matrices of the sensor on the three coordinate axes x, y, and z respectively; Step S13: Linearly transform the GASF matrix values to the interval [0, 1], and then adjust each element value in the matrix to the range of 0 - 255; Step S14: Generate a pseudo deBrujin sequence; Step S15: Assume that the time series length is N, including K time steps. Then, according to the generated deBrujin sequence, select P consecutive time steps of GASF matrices for matrix splicing to obtain the final GASF matrix, where P < K and P can be divisible by K. Finally, convert the spliced GASF matrix into a grayscale image.

3. The method according to claim 1, characterized in that In step S2, using the convolutional neural network for feature extraction includes: Convolution module 1: For the input single-channel image ,in, represents a real number set, H and W represent the height and width of the image respectively. The three layers of convolution in the convolution module 1 extract the 32-dimensional features of the input image in turn. The size of each convolution kernel is , with a step size of 1 and a padding of 1, expressed as: ; in, Indicates Features extracted by layer convolution, and Respectively The weights and biases of the layer convolution, * indicates a two-dimensional convolution operation, is the ReLU activation function, and a 2×2 maximum pooling operation is performed after the third convolution layer to halve the spatial resolution: , Represents the maximum pooling operation, and the output feature map size becomes ; Convolutional module 2: Input , the three-layer convolution of convolution module 2 expands the number of feature channels to 64, and the convolution operation is expressed as: ; in, and Respectively The weights and biases of the layer convolution, , after the sixth convolution layer, a 2×2 maximum pooling is performed to further reduce the resolution: , the output feature map size becomes ; Convolutional module 3: Input , the three-layer convolution of convolution module 3 expands the number of feature channels to 128, and the convolution operation is expressed as: ; in, and Respectively The weights and biases of the layer convolution, , after the ninth convolution layer, a 2×2 maximum pooling is performed: , the output feature map size is ; Fully connected layer part: After convolution and pooling operations, the feature map is flattened into a vector: , the output feature map size is , and input the obtained output into the fully connected layer: , ; in, , , represents the weight of the fully connected layer, and finally outputs a 256-dimensional feature vector: .

4. The method according to claim 1, characterized in that The step S5 specifically includes: First, the channel exchange feature And weight fusion features Normalize: , ; The normalized sequence is then input into the MLP using the linear layer Perform a linear mapping to map the sequence to a higher dimension P: ; ; ; , Further features are extracted through one-dimensional convolution and activation function: ; ; Then based on , Generate space state parameters: ; ; in, , is the state dynamic matrix, , is the input control matrix, , To output the projection matrix, use SSM , Perform dynamic modeling and get output , : ; ; Using gating factors Perform weighted activation: ; ; Fuse the activated modal features and add a residual term: ; Get the final fusion feature .

5. The method according to claim 1, characterized in that The step S6 includes: Use a binary support vector machine for identity legal verification. After obtaining the output features of step S5, initialize the support vector machine using the Gaussian kernel function. The Gaussian kernel function is: ; represents the sample points in the input space, It is a hyperparameter. Through the Gaussian kernel function binary support vector machine, the optimal hyperplane can be found in the high-dimensional space, so as to effectively judge the identity. If the identity is legal, the device can continue to be used, otherwise the device is prohibited from being used.

Citation Information

Patent Citations

  • Smart phone security protection system and method based on Internet

    CN119172755A