Internet of things identity continuous verification method based on multi-mode technology

By continuously collecting and analyzing user's facial images and keyboard input mode data on IoT devices, and using deep learning for cross-modal correlation learning, the problem of unsustainable identity verification in the prior art is solved, and security and user experience are improved.

CN120046133AInactive Publication Date: 2025-05-27STATE GRID HENAN INFORMATION & TELECOMM CO
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510056777.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing IoT authentication method based on multimodal technology does not conduct continuous verification after the user logs in for the first time, which poses a security risk. At the same time, the traditional continuous verification method will reduce the user experience.

Method used

By continuously collecting facial image data and keyboard input mode data when users interact with IoT devices, deep learning technology is used to extract facial semantic features and keyboard usage behavior mode features, and perform cross-modal association learning interaction fusion to achieve continuous verification of user identity.

Benefits of technology

Continuous monitoring and verification of user identities is realized, the security of IoT devices is improved, and the user experience decline caused by traditional continuous verification methods is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046133A_ABST
    Figure CN120046133A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of identity continuous verification, and particularly discloses an Internet of Things identity continuous verification method based on a multi-mode technology, which comprises the following steps: continuously collecting face image data and keyboard input mode data of a user in an interaction process of the user and Internet of Things equipment; the deep learning technology is adopted to process the face image data and the keyboard input mode data of the user to extract the face semantic features and the keyboard use behavior mode features of the user, and then cross-modal association learning interaction fusion is carried out on the face features and the keyboard use behavior mode features of the user, so that the user experience is improved. Therefore, the potential association between the two is captured, and the continuous verification of the user identity is realized. Through the mode, the behavior of the user can be continuously monitored, and an alarm is given out in time when the identity of the user is abnormal, so that the safety of the Internet of Things equipment is effectively improved, and meanwhile, the problem that the user experience is reduced possibly caused by a traditional continuous verification method is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of identity continuous verification, and more specifically, to an Internet of Things identity continuous verification method based on multimodal technology. Background Art

[0002] With the rapid development of Internet of Things (IoT) technology, more and more intelligent devices are connected to the network, providing convenient services for users. However, the security issues of IoT devices are becoming increasingly prominent, especially in user identity verification. Traditional identity verification methods, such as usernames and passwords, have the risk of being cracked or stolen and cannot meet the requirements of high-security identity verification for IoT devices.

[0003] To address this challenge, researchers have begun to explore biometric-based identity verification methods, such as voice recognition and face recognition, which achieve accurate identification of user identities by extracting users' biometric information. Currently, face recognition technology has become a widely used identity verification method in IoT devices due to its contactless, intuitive, and user-friendly features. However, although face recognition has advantages such as contactless and convenience, its recognition accuracy may be affected in complex environments such as changes in lighting and the presence of occlusions. Similarly, to overcome the limitations of single-modal identity verification methods, researchers have proposed multimodal technology-based identity verification methods, which collect multiple biometric information of users, such as face images, fingerprints, voices, etc., to improve the accuracy and reliability of identity verification.

[0004] However, most of the existing multimodal technology-based identity verification methods perform identity verification when users log in for the first time. Once users log in successfully, continuous identity verification is no longer performed. Once a user's account is controlled by a malicious attacker, the attacker can perform illegal operations without being detected, posing certain security risks. At the same time, traditional continuous verification methods may require users to cooperate with additional verification operations regularly, greatly reducing the user experience.

[0005] Therefore, an Internet of Things identity continuous verification method based on multimodal technology is expected. Summary of the Invention

[0006] To solve the above technical problems, this application is proposed. An embodiment of this application provides an Internet of Things identity continuous verification method based on multimodal technology. During the interaction between the user and the Internet of Things device, it continuously collects the user's facial image data and keyboard input pattern data, and uses deep learning technology to process the user's facial image data and keyboard input pattern data to extract the user's facial semantic features and keyboard usage behavior pattern features. Furthermore, by performing cross-modal correlation learning and interactive fusion on the user's facial features and keyboard usage behavior pattern features, it captures the potential correlation between the two to achieve continuous verification of the user's identity. In this way, the user's behavior can be continuously monitored, and an alarm can be issued in a timely manner when there is an abnormality in the user's identity, thereby effectively improving the security of the Internet of Things device and avoiding the problem of the possible decline in user experience caused by traditional continuous verification methods.

[0007] According to one aspect of this application, there is provided an Internet of Things identity continuous verification method based on multimodal technology, which includes:

[0008] Obtain the first-modal data for identity verification and the second-modal data for identity verification, where the first-modal data for identity verification and the second-modal data for identity verification are a facial image and keyboard input pattern data respectively;

[0009] Extract facial features from the first-modal data for identity verification to obtain a facial feature semantic coding feature map;

[0010] Extract keyboard input pattern features from the second-modal data for identity verification to obtain a keyboard input pattern time series coding feature vector;

[0011] Perform attention joint based on fine-grained ablation on the facial feature semantic coding feature map and the keyboard input pattern time series coding feature vector to obtain a facial-keyboard input pattern cross-modal ablation coding feature map;

[0012] Based on the facial-keyboard input pattern cross-modal ablation coding feature map, determine whether the target user is an authorized object.

[0013] Compared with the prior art, the IoT identity continuous verification method based on multimodal technology provided by the present application continuously collects the user's facial image data and keyboard input pattern data during the interaction between the user and the IoT device, and uses deep learning technology to process the user's facial image data and keyboard input pattern data to extract the user's facial semantic features and keyboard usage behavior pattern features. Furthermore, by performing cross-modal correlation learning and interactive fusion on the user's facial features and keyboard usage behavior pattern features, the potential correlation between the two is captured to achieve continuous verification of the user's identity. In this way, the user's behavior can be continuously monitored, and an alarm can be issued in a timely manner when there is an abnormality in the user's identity, thereby effectively improving the security of the IoT device and avoiding the problem of the decline in user experience that may be brought about by traditional continuous verification methods. Description of the Drawings

[0014] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0015] Figure 1 It is a flowchart of the IoT identity continuous verification method based on multimodal technology according to an embodiment of the present application.

[0016] Figure 2 It is a schematic diagram of data flow of the IoT identity continuous verification method based on multimodal technology according to an embodiment of the present application.

[0017] Figure 3 It is a flowchart of sub-step S4 of the IoT identity continuous verification method based on multimodal technology according to an embodiment of the present application.

[0018] Figure 4 It is a flowchart of sub-step S41 of the IoT identity continuous verification method based on multimodal technology according to an embodiment of the present application. Detailed Description of the Embodiments

[0019] As shown in the present application and the claims, unless the context clearly indicates an exception, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "including" and "comprising" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0020] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules may be used and run on a user terminal and / or a server. The modules are merely illustrative, and different aspects of the system and method may use different modules.

[0021] Flowcharts are used in the present application to illustrate the operations performed by the system according to embodiments of the present application. It should be understood that the operations above or below do not necessarily have to be executed precisely in sequence. On the contrary, various steps may be processed in reverse order or simultaneously as needed. At the same time, other operations may also be added to these processes, or one or several operations may be removed from these processes.

[0022] Hereinafter, exemplary embodiments of the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0023] It is worth noting that in the present application, all actions of obtaining data are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the data is located and obtaining authorization from the owner of the corresponding device.

[0024] In view of the technical problems described in the above background art, the present application proposes an Internet of Things identity continuous verification method based on multimodal technology. During the interaction between a user and an Internet of Things device, the method continuously collects the user's facial image data and keyboard input pattern data, and uses deep learning technology to process the user's facial image data and keyboard input pattern data to extract the user's facial semantic features and keyboard usage behavior pattern features. Furthermore, by performing cross-modal correlation learning and interactive fusion on the user's facial features and keyboard usage behavior pattern features, the potential correlation between the two is captured to achieve continuous verification of the user's identity. In this way, the user's behavior can be continuously monitored, and an alarm can be issued in a timely manner when there is an abnormality in the user's identity, thereby effectively improving the security of Internet of Things devices and avoiding the problem of the decline in user experience that may be brought about by traditional continuous verification methods.

[0025] Figure 1 It is a flowchart of an Internet of Things identity continuous verification method based on multimodal technology according to an embodiment of the present application. Figure 2 It is a schematic diagram of data flow of an Internet of Things identity continuous verification method based on multimodal technology according to an embodiment of the present application. As Figure 1 and Figure 2As shown, the method for continuous authentication of Internet of Things identities based on multimodal technology includes the steps of: S1, obtaining first-modal data for authentication and second-modal data for authentication, where the first-modal data for authentication and the second-modal data for authentication are a facial image and keyboard input pattern data, respectively; S2, performing facial feature extraction on the first-modal data for authentication to obtain a facial feature semantic encoding feature map; S3, performing keyboard input pattern feature extraction on the second-modal data for authentication to obtain a keyboard input pattern temporal encoding feature vector; S4, performing attention joint based on fine-grained ablation on the facial feature semantic encoding feature map and the keyboard input pattern temporal encoding feature vector to obtain a facial-keyboard input pattern cross-modal ablation encoding feature map; S5, determining whether the target user is an authorized object based on the facial-keyboard input pattern cross-modal ablation encoding feature map.

[0026] In the above method for continuous authentication of Internet of Things identities based on multimodal technology, in step S1, first-modal data for authentication and second-modal data for authentication are obtained, where the first-modal data for authentication and the second-modal data for authentication are a facial image and keyboard input pattern data, respectively. It should be understood that the facial image reflects the physiological characteristics of an individual, with a relatively fixed facial structure and feature points, while the keyboard input pattern reflects the operation habits formed by the user over a long time, including key pressing force, rhythm, character combination preference, etc. By obtaining facial image data and keyboard input pattern data of different modalities, the user identity can be characterized from multiple dimensions of the user's physiological characteristics and behavioral characteristics, improving the accuracy and reliability of authentication. At the same time, the method for continuous authentication of identities based on the user's facial features and keyboard usage behavior pattern features can perform identity authentication during the natural interaction process between the user and the Internet of Things device, avoiding the cumbersome process of the user actively cooperating for additional verification operations and improving the user experience. In specific implementation, a camera installed on the Internet of Things device can be used to capture a clear facial image in real time during the user's operation; at the same time, the software built into the device is used to record each input operation of the user on the keyboard, including keyboard input pattern data such as the character code corresponding to the key, input timestamp, key pressing force, etc.

[0027] Specifically, when capturing a facial image as the first-modal data, it is first necessary to select a camera suitable for the Internet of Things environment. Considering factors such as cost, power consumption, and size, an embedded camera or a dedicated security camera is an ideal choice. These cameras usually have high-definition resolution and can work under low-light conditions to ensure that clear facial images can be captured in different lighting environments. The camera should be deployed in a place where the user's front view can be clearly captured, while avoiding problems such as reflection or shadow caused by direct light sources.

[0028] To ensure the quality of the captured facial images, the capture timing needs to be reasonably set. On the one hand, automatically triggering the capture process when the user logs in or performs sensitive operations can ensure that the obtained data matches the actual usage scenario; on the other hand, it can also allow the user to actively trigger the image capture to meet different usage requirements. In addition, sensors (such as infrared sensors) can be used to detect the presence of the user and start the capture program at an appropriate time.

[0029] The captured original images may contain unnecessary background information or be affected by environmental factors. Therefore, it is necessary to preprocess the facial images before further processing. This includes steps such as resizing the image, color correction, and denoising, with the aim of improving the effect of the subsequent feature extraction stage. Additionally, if there are multiple users appearing in front of the camera at the same time, face detection algorithms need to be used to locate and crop the facial regions of individual users.

[0030] Since facial images involve personal privacy, relevant laws and regulations must be followed during the collection process, and necessary encryption measures should be taken to protect the security of the data. For example, after preliminary processing of the images on the local device, only the necessary feature point information can be uploaded to the server side, thus reducing the risk of leakage of sensitive information.

[0031] In addition to visual biometric features, the typing behavior of users is also a unique biometric feature. To implement an easy-to-implement and effective keyboard input pattern data acquisition scheme, a lightweight event listener can be integrated into the user interface. This listener captures the timestamp, key value, and the time intervals of key presses and releases for each key event based on existing Web technologies (such as JavaScript) or native application development tools. For Web applications, the KeyboardEvent interface provided by HTML5 allows developers to easily access this information; while in mobile applications, the corresponding operating system API can be used to achieve the same function. This listening mechanism does not require modification of the underlying operating system and is easy to integrate into existing platforms, but attention needs to be paid to cross-browser compatibility issues (for Web applications).

[0032] Next, it is very important to determine the reasonable sampling timing without affecting the user experience. The ideal capture moment should be when performing sensitive operations or when identity verification is necessary, such as on the login screen, password change page, etc. In addition, the capture process can also be triggered when the user reactivates the session after a long period of inactivity. This not only reduces the interference to daily use but also improves user acceptance. However, it is necessary to avoid requesting the user to cooperate with additional identity verification actions too frequently to avoid unnecessary disturbance.

[0033] Extracting keyboard input pattern data that can represent a user's typing habit from the captured keyboard events is one of the key steps. For example, recording the time difference between two adjacent keys to reflect the user's typing speed and rhythm (i.e., key press delay), measuring the duration each key is pressed to reflect the strength and habit of different people (key travel time), counting the usage frequency of specific key combinations (such as Shift+C) because everyone's habits are different (combination key usage frequency), and calculating the ratio of the number of misclicks to the total number of key presses to distinguish novice users from proficient users (error rate). Through the data in the above several dimensions, a preliminary user behavior model can be constructed.

[0034] To ensure that the collected data is properly processed, the original data should first be anonymized to remove any information that can be directly associated with personal identity, and then these data are encrypted and stored or only the necessary feature vectors are transmitted to the server for analysis to minimize potential risks. Additionally, most of the computing work can be considered to be completed on the local device, and only the final results are uploaded to further enhance security. This not only strengthens the user privacy protection measures but also meets modern data security standards, although this may require optimizing the front-end computing performance.

[0035] In the above method for continuous authentication of Internet of Things identities based on multimodal technology, in step S2, facial feature extraction is performed on the first-modal data of the identity authentication to obtain a facial feature semantic encoding feature map. In a specific example of the present application, step S2 includes: inputting the first-modal data of the identity authentication into a facial feature extractor based on the Mobile-Former model to obtain the facial feature semantic encoding feature map. It should be understood that since the facial image contains rich visual information, in order to extract discriminative facial feature representations from the facial image, the present application further uses the Mobile-Former model, which performs excellently in the field of image recognition, to extract features from the facial image. Those of ordinary skill in the art should know that the Mobile-Former model combines the advantages of mobile lightweight models and the Transformer model based on the attention mechanism, and can effectively extract key features in facial images while ensuring a certain computational efficiency. Specifically, the Mobile-Former model adopts a parallel architecture design. The Mobile branch uses a lightweight convolutional neural network to quickly extract local features of the facial image, such as texture, edge and other information, while the Former branch is based on the self-attention mechanism of the Transformer architecture to model the global features of the image, paying attention to the semantic associations and importance weight distributions between different local regions of the face. Through bidirectional information interaction, the local and global features are fused and optimized to effectively highlight the key regions and semantic features closely related to identity recognition in the facial image, such as the feature information of parts such as eyes, nose, and mouth, thereby generating a facial feature semantic encoding feature map.

[0036] In the above-mentioned Internet of Things identity continuous verification method based on multimodal technology, in step S3, keyboard input mode feature extraction is performed on the second-modal data of identity verification to obtain a keyboard input mode time-series encoded feature vector. In a specific example of the present application, step S3 includes: inputting the second-modal data of identity verification into a keyboard input mode feature extractor based on an LSTM model to obtain the keyboard input mode time-series encoded feature vector. It should be understood that the present application takes into account that keyboard input is a behavioral process with time-series characteristics, and static original keyboard input data combinations are difficult to directly use for identity recognition. Therefore, in order to effectively extract the hidden user behavior feature patterns in the keyboard input mode data, the present application further uses an LSTM model to construct a keyboard input mode feature extractor to perform time-series analysis on the second-modal data of identity verification to obtain a keyboard input mode time-series encoded feature vector. Those of ordinary skill in the art should know that due to its unique memory unit and gating mechanism, the LSTM model can effectively handle long-term dependencies in time-series data. In the technical solution of the present application, first, the keyboard input mode data is sorted according to the time order of user input and input into the LSTM model for sequence analysis at each time step. The LSTM model controls the flow and update of information through its internal input gate, forget gate, and output gate, so as to be able to remember key behavior habit information such as the order of user input, input speed, and input rhythm, and generate the corresponding keyboard input mode time-series encoded feature vector, thereby summarizing the unique pattern of the user's keyboard input behavior and providing more easily processable and analyzable feature data for subsequent identity verification.

[0037] In the above-mentioned Internet of Things identity continuous verification method based on multimodal technology, in step S4, attention joint based on fine-grained ablation is performed on the facial feature semantic encoded feature map and the keyboard input mode time-series encoded feature vector to obtain a face-keyboard input mode cross-modal ablation encoded feature map. It should be understood that the present application takes into account that facial features and keyboard input mode features belong to different modalities, coming from the visual domain and the behavioral domain respectively, and there are significant differences in their data structures and feature distributions. Simple splicing or direct fusion cannot fully explore the potential associations and complementary information between the two. Based on this, the present application proposes an attention joint encoding method based on fine-grained ablation, which performs fine-grained correlation analysis on the facial feature semantic encoded feature map and the keyboard input mode time-series encoded feature vector to adaptively focus on the key association regions and feature dimensions between the two, eliminate or weaken the influence of irrelevant or interfering features, and then achieve deep fusion and interaction of cross-modal features, generating more discriminative cross-modal fusion features. Among them, Figure 3 is a flowchart of sub-step S4 of the Internet of Things identity continuous verification method based on multimodal technology according to an embodiment of the present application. As Figure 3As shown, step S4 includes steps: S41, performing feature association strength ablation measurement based on local channels on the keyboard input mode timing encoding feature vector and the facial feature semantic encoding feature map to obtain a set of face-keyboard input mode fine-grained ablation factors; S42, based on the set of face-keyboard input mode fine-grained ablation factors, performing fine-grained ablation modulation on the facial feature semantic encoding feature map to obtain the face-keyboard input mode cross-modal ablation encoding feature map.

[0038] Figure 4 It is a flowchart of sub-step S41 of the Internet of Things identity continuous verification method based on multi-modal technology according to an embodiment of the present application. As Figure 4 shown, step S41 includes steps: S411, performing feature fine-grained decoupling on the facial feature semantic encoding feature map along the channel dimension to obtain a set of facial feature semantic encoding feature matrices; S412, performing cross-domain query interaction based on the attention mechanism on each facial feature semantic encoding feature matrix in the set of the keyboard input mode timing encoding feature vector and the facial feature semantic encoding feature matrices to obtain a set of face-keyboard input mode cross-domain query interaction feature vectors; S413, respectively inputting each face-keyboard input mode cross-domain query interaction feature vector in the set of face-keyboard input mode cross-domain query interaction feature vectors into an ablation measurement function to obtain the set of face-keyboard input mode fine-grained ablation factors.

[0039] More specifically, step S411 is expressed by the formula:

[0040] Decompose(M)={M 1 ,M 2 ,...,M n}

[0041] where V 1 represents the keyboard input mode timing encoding feature vector, M represents the facial feature semantic encoding feature map, M 1 , M 2 , M i and M n respectively represent the 1st, 2nd, ith, and nth facial feature semantic encoding feature matrices in the set of facial feature semantic encoding feature matrices, and n is the number of channels of the facial feature semantic encoding feature map.

[0042] That is, by performing feature decoupling on the facial feature semantic encoding feature map along the channel dimension, a finer-grained facial feature representation is obtained, providing a more detailed feature basis for subsequent ablation analysis.

[0043] More specifically, step S412 includes: performing a linear transformation on the temporal encoding feature vector of the keyboard input mode to obtain a query vector and a value vector; performing a linear transformation on the semantic encoding feature matrix of the facial features to obtain a key matrix; and inputting the query vector, the value vector, and the key matrix into a cross-domain interaction encoder based on a transformer structure to obtain the cross-domain query interaction feature vector of the face-keyboard input mode, which is expressed by the formula:

[0044]

[0045] where, W 1q 、W 1v and W 2k respectively represent a query embedding matrix, a value embedding matrix, and a key embedding matrix, b 1q 、b 1v and b 2k respectively represent different bias terms, represents matrix multiplication operation, V 1q 、V 1v and M ki respectively represent a query vector, a value vector, and the key matrix corresponding to the M i , softmax(·) is a normalized exponential function, (·) T represents the transpose of a vector, h i represents the cross-domain query interaction feature vector between the V 1 and the M i .

[0046] That is, a query vector and a value vector are constructed based on the temporal encoding feature vector of the keyboard input mode, a key matrix is constructed based on each facial feature semantic encoding feature matrix in the set of the facial feature semantic encoding feature matrices, and the correlation between the facial features and the keyboard input mode features is captured through cross-domain query interaction based on a transformer structure. In this process, the transformer structure can effectively process information interaction between different domains, strengthen the significant correlation features between the facial features and the keyboard input mode features through the attention mechanism, and at the same time suppress irrelevant information interference, so as to obtain the correlation interaction feature representation between the facial features and the keyboard input mode features.

[0047] More specifically, step S413 is expressed by the formula:

[0048]

[0049] where, e(·) represents an ablation metric function, max(·) represents a function to take the maximum value, μ i and σ i respectively represent the h iThe characteristic mean and characteristic variance, λ represents the regularization term, e i represents the h i corresponding fine-grained ablation factor of the face-keyboard input mode.

[0050] Specifically, the present application further introduces an ablation metric function to evaluate and quantify the associative interaction feature representations between the various local face features and keyboard input mode features after fine-grained decoupling. Here, the ablation metric function is similar to the ablation analysis in experimental design, used to judge the impact on the final interaction effect after removing a certain specific relationship, helping to identify the degree of influence of the keyboard input mode feature information on different local face features, so as to adaptively adjust the weight distribution of each feature dimension of the face features. For example, if the keyboard input mode feature shows that the user frequently taps the keyboard with a large force, it may mean that the user is in a tense or concentrated state. At this time, the eye and mouth areas in the face features may show corresponding tense or concentrated expression features, and then the association between these face feature areas and the keyboard input mode feature can be focused on.

[0051] Specifically, the step S42 includes: First, input the set of the fine-grained ablation factors of the face-keyboard input mode into an ablation effect encoding module including a normalization function and a masking function to obtain a set of fine-grained ablation weight factors of the face-keyboard input mode, which is expressed by the formula:

[0052]

[0053] where exp(·) represents the exponential function with base e, a i represents the e i corresponding normalized fine-grained ablation factor of the face-keyboard input mode, mask(·) is the masking function, θ is the gating mask threshold, w i is the M i corresponding fine-grained ablation weight factor of the face-keyboard input mode, and M(·) is the ablation effect encoding module including the normalization function and the masking function.

[0054] That is, in order to standardize the weights and exclude irrelevant terms, the present application further performs normalization and masking processing on the generated set of fine-grained ablation factors of the face-keyboard input mode. Here, the normalization function is used to convert all ablation factors into a standard range (such as between 0 and 1) for direct comparison with each other; while the masking function is used to increase the ablation factors of the face features that have a significant association with the keyboard input mode features, and at the same time reduce the ablation factors of the face features that are not relevant or have a low correlation with the keyboard input mode features, so as to concentrate resources and attention on the face features that are most closely related to the keyboard input mode features.

[0055] Then, based on the set of fine-grained ablation weight factors of the face-keyboard input mode, perform weighted modulation on the set of face feature semantic encoding feature matrices to obtain a set of cross-modal ablation encoding feature matrices of the face-keyboard input mode; finally, perform feature aggregation along the channel dimension on the set of cross-modal ablation encoding feature matrices of the face-keyboard input mode to obtain the cross-modal ablation encoding feature map of the face-keyboard input mode, which is expressed by the formula:

[0056] F 1-2 ={M 1 ·w 1 ,M 2 ·w 2 ,...,M n ·w n}

[0057] where w 1 , w 2 and w n are the fine-grained ablation weight factors of the face-keyboard input mode corresponding to the M 1 , the M 2 and the M n respectively, and F 1-2 represents the cross-modal ablation encoding feature map of the face-keyboard input mode.

[0058] That is, based on the generated set of fine-grained ablation weight factors of the face-keyboard input mode, perform fine-grained ablation modulation and feature aggregation along the channel dimension on the original set of face feature semantic encoding feature matrices, so as to strengthen the face features highly related to the keyboard input mode features, thereby generating an optimized and enhanced cross-modal ablation encoding feature representation of the face-keyboard input mode. In this way, not only the key information from the visual domain and the behavioral domain is fused, but also through fine-grained ablation analysis, the focus is more on the key associations between the two, effectively improving the effectiveness and discriminability of the features.

[0059] In the above-mentioned Internet of Things identity continuous verification method based on multimodal technology, in step S5, based on the cross-modal ablation coding feature map of the face-keyboard input mode, it is determined whether the target user is an authorized object. In a specific example of this application, step S5 includes: inputting the cross-modal ablation coding feature map of the face-keyboard input mode into an identity verification module based on a classifier to obtain an identity verification result, and the identity verification result is used to indicate whether the target user is an authorized object. Specifically, the classifier is trained based on known user data, including facial images of authorized users, a large amount of historical keyboard input mode data, and corresponding identity verification label information, to learn and identify the unique facial features and keyboard input behavior patterns of authorized users. During the identity verification process, the classifier based on the learned authorized user feature patterns performs feature recognition and classification on the received cross-modal ablation coding feature map of the face-keyboard input mode, and represents whether the target user is an authorized user in the form of probability or confidence, so as to complete the identity verification accurately and quickly, which not only improves security but also optimizes the user experience.

[0060] More specifically, inputting the cross-modal ablation coding feature map of the face-keyboard input mode into an identity verification module based on a classifier to obtain an identity verification result, and the identity verification result is used to indicate whether the target user is an authorized object, includes: expanding the cross-modal ablation coding feature map of the face-keyboard input mode into a cross-modal ablation coding feature vector of the face-keyboard input mode; using the fully connected layer of the identity verification module to perform fully connected coding on the cross-modal ablation coding feature vector of the face-keyboard input mode to obtain a cross-modal ablation fully connected coding feature vector of the face-keyboard input mode; inputting the cross-modal ablation fully connected coding feature vector of the face-keyboard input mode into the Softmax classification function of the identity verification module to obtain the probability values of the cross-modal ablation coding feature vector of the face-keyboard input mode belonging to each type label, where the type label includes that the target user is an authorized object or the target user is not an authorized object; determining the type label corresponding to the largest probability value as the specified result.

[0061] In a preferred example of this application, inputting the cross-modal ablation coding feature map of the face-keyboard input mode through an identity verification module based on a classifier to obtain an identity verification result includes:

[0062] First, determine the cross-modal ablation coding distribution probability value of the face-keyboard input mode obtained by inputting the cross-modal ablation coding feature map of the face-keyboard input mode through an identity verification module based on a classifier, and the cross-modal ablation coding distribution probability value represents the probability value that the target user is an authorized object;

[0063] Secondly, perform maximum normalization on the cross-modal ablation encoded feature map of the face-keyboard input mode to obtain the cross-modal ablation encoded probability feature map of the face-keyboard input mode, and calculate the power function of each eigenvalue of the cross-modal ablation encoded probability feature map of the face-keyboard input mode with the exponent of one minus the cross-modal ablation encoded distribution probability value of the face-keyboard input mode to obtain the first cross-modal ablation encoded convergence feature map of the face-keyboard input mode, which is expressed by the formula:

[0064] F 1 =F ⊙(1-p)

[0065] where F 1 represents the first cross-modal ablation encoded convergence feature map of the face-keyboard input mode, F represents the cross-modal ablation encoded probability feature map of the face-keyboard input mode, (·) ⊙(1-p) represents calculating the power function of each eigenvalue in the feature map with the exponent of one minus the cross-modal ablation encoded distribution probability value of the face-keyboard input mode, and p represents the cross-modal ablation encoded distribution probability value of the face-keyboard input mode;

[0066] Next, calculate the power function of each eigenvalue of the point difference feature map between the unit feature map and the cross-modal ablation encoded probability feature map of the face-keyboard input mode with the exponent of the cross-modal ablation encoded distribution probability value of the face-keyboard input mode to obtain the second cross-modal ablation encoded convergence feature map of the face-keyboard input mode, which is expressed by the formula:

[0067]

[0068] where F 2 represents the second cross-modal ablation encoded convergence feature map of the face-keyboard input mode, F I represents the unit feature map, represents the point difference, (·) ⊙p represents calculating the power function of each eigenvalue in the feature map with the exponent of the cross-modal ablation encoded distribution probability value of the face-keyboard input mode;

[0069] Then, perform a dot product on the difference between the cross-modal ablation encoded probability feature map of the face-keyboard input mode and one minus the cross-modal ablation encoded distribution probability value of the face-keyboard input mode to obtain the first cross-modal ablation encoded limited feature map of the face-keyboard input mode, which is expressed by the formula:

[0070] F 3 =F⊙(1 - p)

[0071] where F 3 represents the first cross-modal ablation encoded limited feature map of the face-keyboard input mode, and ⊙ represents the dot product;

[0072] Next, multiply the point difference feature map with the cross-modal ablation coding distribution probability value of the face-keyboard input mode to obtain a second cross-modal ablation coding limited feature map of the face-keyboard input mode, which is expressed by the formula:

[0073]

[0074] where F 4 represents the second cross-modal ablation coding limited feature map of the face-keyboard input mode;

[0075] Then, after multiplying the first cross-modal ablation coding convergence feature map of the face-keyboard input mode and the second cross-modal ablation coding convergence feature map of the face-keyboard input mode, perform a dot addition with the first cross-modal ablation coding limited feature map of the face-keyboard input mode and the second cross-modal ablation coding limited feature map of the face-keyboard input mode to obtain an optimized cross-modal ablation coding feature map of the face-keyboard input mode, which is expressed by the formula:

[0076]

[0077] where F' represents the optimized cross-modal ablation coding feature map of the face-keyboard input mode, represents vector dot addition, represents dot addition;

[0078] Finally, input the optimized cross-modal ablation coding feature map of the face-keyboard input mode into the classifier-based authentication module to obtain an authentication result.

[0079] Here, considering that the face feature semantic coding feature map and the keyboard input mode time series coding feature vector respectively represent the semantic coding features of the target object's face image and the keyboard input time series mode coding features of the target object, during the cross-domain attention joint coding based on fine-grained ablation, due to the fine-grained ablation differences caused by the time series distribution differences and modal differences of the data sources, resulting in cross-domain fine-grained semantic joint coding differences, the cross-modal ablation coding feature map of the face-keyboard input mode will have probability convergence divergence based on different feature coding interactions, thereby affecting the accuracy of the authentication result obtained through the classifier-based authentication module.

[0080] Based on this, by taking the cross-entropy formal power series of the regression analysis probability of the cross-modal ablation coding feature map in the face-keyboard input mode as the probability distribution convergence limit, the class probability convergence approximation of the cross-modal ablation coding feature map in the face-keyboard input mode is carried out on the basis of the combination of the feature set distribution and the probability density distribution of the cross-modal ablation coding feature map in the face-keyboard input mode, so as to further use the probability distribution cross-entropy of the cross-modal ablation coding feature map in the face-keyboard input mode as the objective function to guide the limit recovery strategy, and realize the common agility of the convergence of the feature set with probability convergence divergence of the cross-modal ablation coding feature map in the face-keyboard input mode to the probability density distribution space, and improve the accuracy of the authentication result obtained by the cross-modal ablation coding feature map in the face-keyboard input mode through the classifier-based authentication module.

[0081] After determining whether the target user is an authorized object, the Internet of Things (IoT) device or system needs to take a series of measures to ensure security and user experience. This not only involves the processing of the verification result, but also includes adjusting the system's response strategy according to the verification result, updating the user status, and providing corresponding feedback information to the user. The following are the specific implementation steps of this process:

[0082] First, the system needs to analyze the authentication result. Usually, the authentication module will output a confidence score or directly give a binary classification result (i.e., "the target user is an authorized object" or "the target user is not an authorized object"). To improve accuracy, a threshold can be set, and only when the confidence score exceeds this threshold is the verification considered successful. For a high-confidence pass situation, if the user's feature data highly matches that of the known authorized object and the confidence score far exceeds the set threshold, the access permission can be directly granted. At this time, the system should record this login event, including information such as the timestamp and the device ID used, for subsequent auditing. When the confidence is close to but does not reach the preset threshold, the system can trigger additional security checks, such as asking the user to provide a one-time verification code (OTP), or making a more stringent biometric verification request (such as fingerprint scanning). This approach can increase the difficulty for potential attackers while not affecting most legitimate users. If the verification result shows that the user is not an authorized object, the system should immediately block the access and may trigger an alarm mechanism to notify the administrator. At the same time, a clear error message should be provided to the user, informing the reason and the next operation guide (such as retrying, contacting the support staff, etc.).

[0083] Regardless of the verification result, the user's status information should be updated in a timely manner. For a successful verification, update information such as the user's online status, the most recent login time, and IP address; while for a failed case, mark it as an abnormal login attempt and increment the corresponding counter. These information are crucial for monitoring account activity patterns and detecting potential threats. In addition, all processes involving authentication should be completely recorded to form a detailed log file. These logs should not only contain basic login information, but also include a summary of biometric data (encrypted) collected during each verification process, the specific verification steps executed, the final decision, and its reasons. Good log management helps with post - hoc analysis and problem troubleshooting, and also meets compliance requirements.

[0084] The design of the user interface should fully consider how to effectively convey the verification result to the user. For a successful verification, a short success message can be displayed and smoothly transition to the next operation interface. For a failed case, special attention should be paid to avoid revealing too many details, so as not to help attackers optimize their attack strategies. For example, do not explicitly indicate which part of the verification failed, but use a general error message such as "Your identity cannot be confirmed. Please try again later". Additionally, visual and auditory elements can be used to enhance the feedback effect, such as playing a specific sound prompt or changing the screen background color to attract the user's attention. However, it should be noted that any additional interaction should be quick and non - intrusive, ensuring that it does not interrupt the normal usage process. Moreover, considering the support for a multilingual environment, the system should be able to automatically switch languages according to the user's preference to ensure that each user can obtain the most comfortable operation experience.

[0085] Based on the verification result, the system may need to dynamically adjust its security policies. For example, temporarily lock the account for a period of time after multiple consecutive failures, require the user to provide additional identification materials, or restrict certain sensitive operations until more rigorous authentication is completed. These measures aim to balance the relationship between security and convenience, prevent abuse, and minimize the impact on normal users. For accounts in a high - risk state for a long time, more advanced protection mechanisms can be activated, such as forcing a password change, enabling multi - factor authentication (MFA), or even manual review. The implementation of such policies often requires customization based on specific application scenarios and business logics.

[0086] In summary, the IoT identity continuous verification method based on multimodal technology according to the embodiments of the present application is elucidated. During the interaction between the user and the IoT device, it continuously collects the user's facial image data and keyboard input pattern data, and uses deep learning technology to process the user's facial image data and keyboard input pattern data to extract the user's facial semantic features and keyboard usage behavior pattern features. Furthermore, by performing cross-modal correlation learning and interactive fusion on the user's facial features and keyboard usage behavior pattern features, the potential correlation between the two is captured to achieve continuous verification of the user's identity. In this way, the user's behavior can be continuously monitored, and an alarm can be issued in a timely manner when there is an abnormality in the user's identity, thereby effectively improving the security of the IoT device and avoiding the problem of the decline in user experience that may be brought about by traditional continuous verification methods.

[0087] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitating understanding, rather than limitations. The above details do not limit the present invention to necessarily adopt the above specific details for implementation.

[0088] In the above embodiments, the descriptions of each embodiment have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the unit division is only a logical function division, and there can be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0089] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claims involved.

[0090] In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. A plurality of units stated in the system claims can also be implemented by one unit through software or hardware.

[0091] Finally, it should be noted that the above description has been given for purposes of illustration and description. In addition, the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for continuous identity verification of the Internet of Things based on multimodal technology, characterized in that: include: Acquire identity verification first modality data and identity verification second modality data, wherein the identity verification first modality data and the identity verification second modality data are facial image and keyboard input mode data, respectively; Performing facial feature extraction on the identity verification first modality data to obtain a facial feature semantic encoding feature map; Performing keyboard input mode feature extraction on the identity verification second modality data to obtain a keyboard input mode temporal coding feature vector; Performing fine-grained ablation-based attention union on the facial feature semantic encoding feature map and the keyboard input mode temporal encoding feature vector to obtain a face-keyboard input mode cross-modal ablation encoding feature map; Based on the face-keyboard input mode cross-modal ablation coding feature map, it is determined whether the target user is an authorized object.

2. The method for continuous IoT identity verification based on multimodal technology according to claim 1 is characterized in that: Extracting facial features from the identity verification first modality data to obtain a facial feature semantic encoding feature map includes: The identity verification first modality data is input into a facial feature extractor based on a Mobile-Former model to obtain the facial feature semantic encoding feature map.

3. The method for continuous Internet of Things identity verification based on multimodal technology according to claim 2 is characterized in that: Performing keyboard input mode feature extraction on the identity verification second modality data to obtain a keyboard input mode temporal coding feature vector includes: The second modality data of identity authentication is input into a keyboard input pattern feature extractor based on an LSTM model to obtain a temporal encoding feature vector of the keyboard input pattern.

4. The method for continuous identity verification of the Internet of Things based on multimodal technology according to claim 3 is characterized in that: Performing a fine-grained ablation-based attention union on the facial feature semantic encoding feature map and the keyboard input mode temporal encoding feature vector to obtain a face-keyboard input mode cross-modal ablation encoding feature map, including: Performing a local channel-based feature correlation strength ablation measurement on the keyboard input mode temporal encoding feature vector and the facial feature semantic encoding feature map to obtain a set of face-keyboard input mode fine-grained ablation factors; Based on the set of fine-grained ablation factors of the face-keyboard input mode, fine-grained ablation modulation is performed on the facial feature semantic encoding feature map to obtain the face-keyboard input mode cross-modal ablation encoding feature map.

5. The method for continuous Internet of Things identity verification based on multimodal technology according to claim 4 is characterized in that: The keyboard input mode temporal encoding feature vector and the facial feature semantic encoding feature map are subjected to feature association strength ablation measurement based on local channels to obtain a set of fine-grained ablation factors of the facial-keyboard input mode, including: Performing fine-grained feature decoupling along the channel dimension on the facial feature semantic encoding feature map to obtain a set of facial feature semantic encoding feature matrices; Performing cross-domain query interaction based on the attention mechanism on each facial feature semantic encoding feature matrix in the set of the keyboard input mode temporal encoding feature vector and the facial feature semantic encoding feature matrix to obtain a set of face-keyboard input mode cross-domain query interaction feature vectors; Each face-keyboard input mode cross-domain query interaction feature vector in the set of face-keyboard input mode cross-domain query interaction feature vectors is input into the ablation metric function to obtain a set of fine-grained ablation factors of the face-keyboard input mode.

6. The method for continuous Internet of Things identity verification based on multimodal technology according to claim 5 is characterized in that: Performing cross-domain query interaction based on the attention mechanism on each facial feature semantic encoding feature matrix in the set of the keyboard input mode temporal encoding feature vector and the facial feature semantic encoding feature matrix to obtain a set of face-keyboard input mode cross-domain query interaction feature vectors, including: Performing a linear transformation on the keyboard input mode temporal coding feature vector to obtain a query vector and a value vector; Performing a linear transformation on the facial feature semantic encoding feature matrix to obtain a key matrix; The query vector, the value vector and the key matrix are input into a cross-domain interaction encoder based on an imitation transformer structure to obtain the cross-domain query interaction feature vector of the face-keyboard input mode.

7. The method for continuous IoT identity verification based on multimodal technology according to claim 6 is characterized in that: Based on the set of fine-grained ablation factors of the face-keyboard input mode, fine-grained ablation modulation is performed on the facial feature semantic encoding feature map to obtain the face-keyboard input mode cross-modal ablation encoding feature map, including: Inputting the set of face-keyboard input mode fine-grained ablation factors into an ablation effect encoding module including a normalization function and a mask function to obtain a set of face-keyboard input mode fine-grained ablation weight factors; Based on the set of face-keyboard input mode fine-grained ablation weight factors, weighted modulate the set of facial feature semantic encoding feature matrices to obtain a set of face-keyboard input mode cross-modal ablation encoding feature matrices; The set of the face-keyboard input mode cross-modal ablation coding feature matrices is subjected to feature aggregation along the channel dimension to obtain the face-keyboard input mode cross-modal ablation coding feature map.

8. The method for continuous Internet of Things identity verification based on multimodal technology according to claim 7 is characterized in that: Determining whether the target user is an authorized object based on the face-keyboard input mode cross-modal ablation coding feature map includes: The face-keyboard input mode cross-modal ablation coding feature map is input into a classifier-based identity authentication module to obtain an identity authentication result, and the identity authentication result is used to indicate whether the target user is an authorized object.

9. The method for continuous IoT identity verification based on multimodal technology according to claim 8, characterized in that: Inputting the face-keyboard input mode cross-modal ablation coding feature map into a classifier-based identity authentication module to obtain an identity authentication result, wherein the identity authentication result is used to indicate whether the target user is an authorized object, including: Expanding the face-keyboard input mode cross-modal ablation coding feature map into a face-keyboard input mode cross-modal ablation coding feature vector; Using the fully connected layer of the identity verification module to fully connect the face-keyboard input mode cross-modal ablation coding feature vector to obtain the face-keyboard input mode cross-modal ablation fully connected coding feature vector; Inputting the face-keyboard input mode cross-modal ablation fully connected encoding feature vector into the Softmax classification function of the identity authentication module to obtain the probability value of the face-keyboard input mode cross-modal ablation encoding feature vector belonging to each type label, wherein the type label includes whether the target user is an authorized object or the target user is not an authorized object; The type label corresponding to the largest probability value among the probability values ​​is determined as the designated result.

Citation Information

Cited By

  • Verification code intelligent identification and interaction method based on multi-modal large model

    CN120470578A

  • A verification code intelligent recognition and interaction method based on multimodal large model

    CN120470578B

  • Continuous identity authentication method and device based on multi-modal user behavior fusion, equipment and medium

    CN120951306A