A method of verifying identity and a computing device

By collecting and analyzing the feature vectors, trajectories, and temporal features of dynamic gesture data, the problems of high cost and environmental sensitivity of AI glasses authentication have been solved, achieving low-cost and high-security authentication.

CN122113069APending Publication Date: 2026-05-29ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2026-02-27
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing AI glasses authentication methods based on iris or voiceprint are costly or affected by the environment, resulting in low recognition accuracy and poor user experience.

Method used

Authentication is performed using dynamic gesture data. By acquiring the feature vectors, trajectories, and temporal features of dynamic gesture videos, and combining this with an inertial measurement unit to eliminate the influence of camera displacement, authentication is achieved.

Benefits of technology

It achieves low-cost, environment-insensitive identity verification, is highly difficult to mimic dynamic gestures, has strong security, and provides a good user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113069A_ABST
    Figure CN122113069A_ABST
Patent Text Reader

Abstract

The specification provides a method for verifying identity, comprising: acquiring dynamic gesture data of a user collected when the user performs identity verification; the dynamic gesture data is determined according to at least one of the following: a feature vector corresponding to a dynamic gesture video, a trajectory of a dynamic gesture, and a timing feature of the trajectory of the dynamic gesture; verifying the identity of the user according to the similarity between a stored dynamic gesture template and the dynamic gesture data; the dynamic gesture template is generated according to dynamic gesture data of the user collected when the user registers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification belong to the field of artificial intelligence technology, and in particular relate to a method and computing device for verifying identity. Background Technology

[0002] Artificial Intelligence (AI) glasses are smart glasses with built-in AI assistants, enabling functions such as voice interaction, real-time translation, and information reminders. In some situations, users need to use AI glasses to perform sensitive actions such as payments. To protect user account security, AI glasses need to verify the user's identity before granting permission to perform the sensitive action.

[0003] In related technologies, AI glasses typically verify user identity based on biometrics such as iris recognition or voiceprints. Iris-based solutions require the addition of expensive optical modules to the AI ​​glasses, resulting in high hardware costs. Voiceprint-based methods, on the other hand, are subject to environmental requirements. For example, in noisy environments, the recognition accuracy is easily affected by ambient noise. Furthermore, voiceprint-based methods require users to read a series of words, which can be inconvenient for users to speak in quiet public places, leading to a poor user experience. Summary of the Invention

[0004] The purpose of this specification is to provide a method and computing device for verifying identity.

[0005] The first aspect of this specification provides a method for verifying identity, including:

[0006] Acquire dynamic gesture data of the user collected during user authentication; the dynamic gesture data is determined according to at least one of the following: feature vector corresponding to the dynamic gesture video, trajectory of the dynamic gesture, and temporal features of the trajectory of the dynamic gesture.

[0007] The user's identity is verified based on the similarity between the stored dynamic gesture template and the dynamic gesture data; the dynamic gesture template is generated based on the user's dynamic gesture data collected during user registration.

[0008] A second aspect of this specification provides an identity verification device, comprising:

[0009] The dynamic gesture data acquisition module is used to acquire the dynamic gesture data of the user collected during the user's authentication process; the dynamic gesture data is determined based on at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal characteristics of the trajectory of the dynamic gesture.

[0010] The identity verification module is used to verify the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data; the dynamic gesture template is generated based on the user's dynamic gesture data collected during user registration.

[0011] A third aspect of this specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the aforementioned method for verifying identity.

[0012] A fourth aspect of this specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, it implements the aforementioned method for verifying identity.

[0013] This specification provides a fifth aspect of a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method for verifying identity.

[0014] This manual describes a method for user authentication based on collected dynamic gesture data. This method is low-cost and unaffected by environmental conditions. Furthermore, the method uses dynamic gestures for verification, which are more difficult to imitate than static gestures, thus offering stronger security. The method provided in this manual offers a convenient and low-cost way to verify user identity. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the trajectory of a dynamic gesture in one embodiment;

[0017] Figure 2 This is a flowchart of a method for verifying identity in one embodiment;

[0018] Figure 3 This is a block diagram of an embodiment of a method for verifying identity. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0020] This specification provides a method for verifying identity. The method provided in this specification relies on collected dynamic gesture data of the user for authentication. Dynamic gestures can include overall hand movements and / or relative movements between different positions of the hand. Overall hand movements can be dynamic actions such as moving the hand from one place to another. Dynamic gestures can be dynamic actions such as changing the hand from a clenched fist to an open fist.

[0021] First, the application scenarios of this manual will be explained. The methods described in this manual can be applied to devices that can capture users' dynamic gestures, such as wearable devices like AI glasses, and smart devices like mobile phones. This manual does not limit the application to these devices.

[0022] To verify user identity, users first need to register a dynamic gesture template, which is generated based on the user's dynamic gesture data collected during registration. The following will explain the process of registering a dynamic gesture template.

[0023] When a user first powers on and sets up the device, or when they select to create a dynamic gesture password in the device's settings interface, the registration process can be initiated, which involves collecting the user's dynamic gesture data.

[0024] The following section will explain several methods for obtaining dynamic gesture data.

[0025] First, dynamic gesture data can be acquired through the camera equipped on the device.

[0026] First, the user's dynamic gesture video can be captured using a first camera.

[0027] Optionally, to capture higher-quality gesture templates, ambient light detection can be performed. Specifically, the device first assesses the current ambient lighting conditions. If the light is too dim or too bright, it will prompt the user to move to a location with suitable lighting via the screen (or, in the case of AI glasses, the display interface shown in front of the user) or voice, to ensure that the camera can clearly capture hand details.

[0028] Optionally, motion space guidance can be provided when capturing dynamic gesture videos. For example, a reference area box can be displayed on the screen, prompting the user to complete the dynamic gesture within the reference area box. When this method is applied to AI glasses, a virtual AR box can be projected onto the screen as the reference area box. This can improve the effectiveness of data collection.

[0029] Secondly, the dynamic gesture data is obtained based on the dynamic gesture video.

[0030] The dynamic gesture data can be determined based on at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal characteristics of the trajectory of the dynamic gesture. These three types of data are referred to as the first data in the following text, and will be explained in detail below.

[0031] The feature vector corresponding to a dynamic gesture video is obtained by extracting features from the video. By extracting features from the dynamic gesture video, it can be converted into a vector. This vector can reflect the characteristics of the user's dynamic gestures, the characteristics of the user's hand, etc.

[0032] Optionally, in order to reduce the influence of the environmental background on the dynamic gesture template, the model can be trained or configured appropriately so that the feature vectors extracted by the model from the dynamic gesture video are more related to the dynamic gesture and less related to the environmental background.

[0033] Specifically, hand recognition can be used to segment hands from videos. Alternatively, a contrastive learning approach can be employed. For a given dynamic gesture video, videos of the same gesture but with different backgrounds are used as positive samples, while videos of other dynamic gestures with similar backgrounds are used as negative samples. This allows the feature extraction model to be trained so that the model extracts more similar features for the same dynamic gesture, while extracting more dissimilar features for different gestures. Of course, these two examples are merely illustrations and do not represent a limitation of this specification.

[0034] The trajectory of a dynamic gesture and its temporal characteristics can be obtained by tracking several key points of the hand in the dynamic gesture video.

[0035] Multiple key points can be preset, such as fingertips, palms, and / or finger joints. The system can identify the positions of these preset key points based on the captured dynamic gesture video and track changes in these positions to obtain the position of the key points in each video frame. For each key point, connecting the positions of the key points in each video frame yields the two-dimensional trajectory of the dynamic gesture. It should be noted that when multiple key points exist, the trajectory of the dynamic gesture can include the trajectory corresponding to each key point. The two-dimensional trajectory of the dynamic gesture can be as follows: Figure 1As shown. By acquiring the trajectory of a dynamic gesture, we can extract a set of features such as the shape, length, and turning points of the trajectory.

[0036] By adding temporal features (i.e., the time of the corresponding video frame in the video) to the trajectory of a dynamic gesture, the temporal characteristics of the dynamic gesture trajectory can be obtained. By acquiring these temporal characteristics, the rhythm of the user's dynamic gestures can be analyzed, including the total time of the movement, and patterns of change in speed and acceleration. These constitute the user's unique force application habits or movement rhythm.

[0037] After acquiring at least one of the above three types of data, in an optional implementation, at least one of the three types of data can be directly used as dynamic gesture data, that is, as a dynamic gesture template. In other words, the extracted first data can be used as the dynamic gesture data, and the first data includes at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal features of the trajectory of the dynamic gesture;

[0038] In another optional implementation, to reduce the workload and storage space required for authentication, if at least two of the three types of data are collected, these two types of data can be fused to obtain the fused features. In other words, the extracted at least two types of first data can be weighted and fused to obtain the dynamic gesture data. Thus, during authentication, it is not necessary to match all three types of data; authentication can be completed by matching only one type of data.

[0039] Second, dynamic gesture data can be acquired through the device's camera and depth camera.

[0040] Specifically, firstly, the user's dynamic gesture video is captured by a first camera, and then a three-dimensional dynamic gesture video is captured by a depth camera.

[0041] A depth camera refers to a camera capable of acquiring stereoscopic data (spatial / 3D information). For example, it could be a Time-of-Flight (ToF) camera, which emits modulated light and measures the round-trip time difference to obtain the distance between the object and the camera in each video frame. Alternatively, a depth camera can be a structured light camera, which emits specific patterns (dots / stripes) and uses pattern deformation to infer the shape of objects in each video. Of course, these examples are not intended to limit the definition of a depth camera.

[0042] A depth camera can capture three-dimensional dynamic gesture videos, which can be used to more accurately determine the three-dimensional trajectory of the dynamic gestures.

[0043] Furthermore, for details on the method of capturing dynamic gesture videos of users using the first camera, please refer to the description of the first method, which will not be repeated here.

[0044] Secondly, the dynamic gesture data can be obtained from the dynamic gesture video.

[0045] The dynamic gesture video is also determined based on the three types of first data mentioned above, and the specific determination method is the same as before, so it will not be repeated here. The method for determining the feature vector corresponding to the dynamic gesture video in the first data is the same as before, so it will not be repeated here.

[0046] Regarding the trajectory and temporal characteristics of dynamic gestures, they can be obtained by tracking several key points of the hand in dynamic gesture videos and 3D dynamic gesture videos. The trajectory of the dynamic gesture is a 3D trajectory. Unlike the previous method, depth information can be obtained from the 3D dynamic gesture video to determine the 3D trajectory.

[0047] When determining the specific three-dimensional trajectory, you can first obtain the two-dimensional trajectory using the method mentioned above, and then obtain the depth information of the points in the two-dimensional trajectory based on the three-dimensional dynamic gesture video, and combine them to obtain the three-dimensional trajectory.

[0048] Additionally, optionally, when a user makes dynamic gestures, slight camera displacement may occur. For example, if this method is applied to AI glasses, the user's head may shake slightly, causing slight displacement of the AI ​​glasses worn by the user. To eliminate the influence of camera displacement, if an inertial measurement unit is installed on the device (such as AI glasses), the user's posture data can be acquired through the inertial measurement unit of the AI ​​glasses; the dynamic gesture data can then be compensated based on the posture data to eliminate the influence of the user's head movement from the dynamic gesture data.

[0049] It should be noted that the dynamic gesture data used for compensation here refers to at least one of the three types of primary data.

[0050] After collecting the dynamic gesture template, the user's dynamic gesture template can be stored. Optionally, to protect the user's privacy and data security, the dynamic gesture template can be stored in an encrypted manner.

[0051] Regarding the storage location of the dynamic gesture templates, in one alternative implementation, they can be stored locally on the device, allowing authentication even when the device is offline. In another alternative implementation, the dynamic gesture templates can be stored on a server. This method is applicable to wearable devices with limited storage space, thus saving device storage space.

[0052] In one optional implementation, not only can the user's dynamic gesture templates be obtained, but also the user's static gesture templates. Multiple dynamic gesture templates can be stored, with different homomorphic gesture templates indexed by different static gestures. Static gestures differ from dynamic gestures in that they refer to gestures that remain stationary, such as the OK sign or a clenched fist.

[0053] In other words, multiple dynamic gesture templates can be pre-stored, and each dynamic gesture template is associated with a corresponding static gesture template, which is generated based on the static gesture images of the user collected during user registration. During verification, the static gesture template can be verified first, and only after the static gesture template is successfully verified will the verification of the dynamic gesture template proceed.

[0054] In one alternative implementation, different dynamic gesture templates can correspond to different sensitive actions. For example, some dynamic gesture templates can be used for authentication of sensitive actions such as unlocking, while others can be used for authentication of sensitive actions such as payment.

[0055] For example, combination one is a static fist gesture + a dynamic S-shaped gesture in the air; this combination one can be used for identity verification during payment. Combination two is a static V-shaped gesture + a dynamic forward-pushing gesture; this combination two can be used for unlocking.

[0056] The following section will explain how to obtain static gesture templates.

[0057] Static gesture templates can be captured using the device's camera. Specifically, a static gesture image of the user can be captured using a first camera. Then, a static gesture template can be obtained based on the static gesture image.

[0058] For capturing static gesture images, they can be captured separately from dynamic gesture videos. For example, the system can first prompt the user to perform an initial static gesture, and then prompt the user to perform a dynamic gesture.

[0059] In an alternative implementation, features of a static gesture image can be extracted as a static gesture template.

[0060] Similar to the features of dynamic gesture videos mentioned above, in order to eliminate the influence of environmental background on static gesture templates, in an optional implementation, hand recognition can be performed on static gesture images, and a separate image corresponding to the hand can be identified, and features can be extracted only from that image.

[0061] In another alternative implementation, the model can be trained so that the features extracted by the model are more relevant to the static gesture itself, and less relevant to the background of the static gesture. Specific training methods are detailed above and will not be repeated here.

[0062] After obtaining the static gesture template, it can be associated and stored with its corresponding dynamic gesture template. The storage method and location are the same as for the dynamic gesture template, and will not be repeated here.

[0063] The following section explains the authentication methods used when users need to perform sensitive actions while using the device. Sensitive actions may include unlocking, viewing stored private data, or making payments.

[0064] like Figure 2 As shown, this specification provides a method for verifying identity, which includes the following steps:

[0065] Step 201: Obtain the user's dynamic gesture data collected during the user's authentication process.

[0066] The dynamic gesture data is determined based on at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal characteristics of the trajectory of the dynamic gesture.

[0067] The method for obtaining dynamic gesture data is the same as the method for obtaining the dynamic gesture data corresponding to the dynamic gesture template during registration, and will not be repeated here.

[0068] Furthermore, in an optional implementation, a static gesture template is pre-stored as described above. In this case, it is also possible to: obtain the user's static gesture data collected during user authentication by using a static gesture image of the user captured by the first camera.

[0069] The method for obtaining static gesture data is the same as the method for obtaining static gesture templates mentioned earlier, and will not be repeated here.

[0070] Step 203: Verify the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data.

[0071] As mentioned above, the dynamic gesture template is generated based on the user's dynamic gesture data collected during user registration.

[0072] In an alternative implementation, if the dynamic gesture template is stored in encrypted form, the encrypted dynamic gesture template can be decrypted first.

[0073] Regarding the specific implementation of step 203, in one optional embodiment, the similarity between the dynamic gesture data and the dynamic gesture template can be directly calculated. If the similarity between the two is greater than a preset threshold, the authentication is determined to be successful; otherwise, the authentication is determined to be unsuccessful.

[0074] If the dynamic gesture template includes multiple types of first data, the similarity between each type of first data and the corresponding data in the dynamic gesture template can be calculated separately. The following will explain the similarity calculation methods for the three types of first data.

[0075] For the feature vectors corresponding to dynamic gesture videos, the similarity between two features can be directly calculated. Similarity can be calculated using methods such as cosine similarity; this specification does not limit the specific methods used.

[0076] For the trajectory of a dynamic gesture, the feature vector corresponding to the trajectory of the dynamic gesture can be extracted, and the similarity between the feature vector and the feature vector of the trajectory in the dynamic gesture template can be calculated.

[0077] Regarding the temporal features of the trajectory of dynamic gestures, in an optional implementation, a method similar to that described above can also be used, that is, first extracting feature vectors and then calculating vector similarity.

[0078] In another alternative implementation, this can be accomplished using Dynamic Time Warping (DTW). DTW is an algorithm for comparing the similarity of two time series of different lengths and speeds by calculating their minimum matching distance by "stretching / compressing" the sequences along the time axis.

[0079] Specifically, the DTW comparison process includes the following steps: Two temporal features in the dynamic gesture template and dynamic gesture data are referred to as the first time series and the second time series, respectively. Then, the following steps are taken: First and second time series are obtained; local distances are calculated for each sampling point in the first time series and each sampling point in the second time series to construct a local distance matrix; based on the local distance matrix, the minimum cumulative distance from the starting position of the matrix to each matrix element is calculated using dynamic programming to obtain a cumulative cost matrix; in the cumulative cost matrix, the alignment path with the minimum cumulative cost is determined by backtracking from the ending position to the starting position; the cumulative cost along the alignment path is used as the DTW distance between the two time series to characterize the similarity between the first and second time series.

[0080] The DTW distance can be used to determine the similarity of two time series features with relatively high accuracy.

[0081] After obtaining the similarity of multiple sets of first data, in one optional implementation, the multiple similarities can be weighted to obtain a final similarity, and the authentication is determined based on whether the final similarity is greater than a preset threshold. In another optional implementation, the similarity of each set of first data can be individually judged to see if it meets the similarity condition (e.g., whether it is greater than the threshold), and the authentication is determined to be successful if all the first data meet the corresponding similarity condition.

[0082] By judging the similarity of multiple primary data, misjudgments caused by accidental bias of a single feature are effectively avoided, and a more accurate identity authentication system is constructed.

[0083] As mentioned above, static gesture templates can also be stored. If static gesture templates corresponding to each dynamic gesture template are stored in advance, a target static gesture template that matches the user's static gesture data can be determined from the stored static gesture templates, and the dynamic gesture template associated with the target static gesture template can be used as the target dynamic gesture template.

[0084] Optionally, static gesture data can be compared with all static gesture templates to determine a target static gesture template that matches the static gesture data. Here, "matching" can refer to the template that is most similar to the static gesture data, and whose similarity to the static gesture data is greater than a preset threshold.

[0085] Optionally, if there are preset verification scenarios corresponding to each static gesture template, the static gesture template corresponding to the current verification scenario (such as payment) can be obtained first. If the static gesture template matches the static gesture data, the static gesture template can be used as the target static gesture template.

[0086] In addition, if the static gesture template is stored in encrypted form, it can be decrypted first to facilitate verification.

[0087] If a target static gesture template that matches the static gesture data cannot be determined, authentication can be considered a failure. This eliminates the need for subsequent matching of dynamic gesture templates, saving computational resources.

[0088] Once a target static gesture template and its corresponding dynamic gesture template are determined, the user's identity is verified based on the similarity between the target dynamic gesture template and the dynamic gesture data.

[0089] Static gesture templates reduce the computational overhead of verification and also increase the difficulty for attackers to imitate them.

[0090] In another alternative implementation, anti-counterfeiting verification can also be performed to defend against advanced spoofing attacks and ensure the authenticity of identity verification.

[0091] First, it can perform liveness detection, such as using artificial intelligence models to analyze whether dynamic gesture videos originate from real hands.

[0092] Second, dynamic tokens can be set. Specifically, to prevent replay attacks, the system will randomly generate a "challenge" (such as "wave upwards first") during verification. The system uses a sequence action recognition model to first determine whether this random "response" is completed correctly, and only then does it perform the core gesture template comparison.

[0093] Specifically, the system can display the target challenge gestures that the user needs to perform; and capture video of the user's challenge gestures through the first camera of the AI ​​glasses.

[0094] It can determine whether the collected challenge gesture matches the target challenge gesture. If they do not match, no further verification is performed.

[0095] Furthermore, in some cases, challenge gesture videos and dynamic gesture videos are captured separately. This makes it easier for the system to distinguish which part is the dynamic gesture and which part is the challenge gesture. In such cases, to avoid replay attacks, continuity detection can be performed between the two videos to determine whether the challenge gesture and the dynamic gesture were filmed at the same location, thereby preventing replay attacks.

[0096] Specifically, if the challenge gesture video matches the target challenge gesture, a continuity detection is performed on the challenge gesture video and the dynamic gesture video to determine whether the challenge gesture video and the dynamic gesture video were filmed at the same location.

[0097] Among them, continuity detection is to determine whether the start frame of one video and the end frame of another video are continuous. Through continuity detection, it can be determined whether two gesture videos were filmed at the same location, thereby avoiding replay attacks.

[0098] If the continuity detection determines that the challenge gesture video and the dynamic gesture video were filmed at the same location, the user's identity can be verified based on the similarity between the stored dynamic gesture template and the dynamic gesture data.

[0099] Third, it can also issue alerts for abnormal behavior.

[0100] Specifically, user behavior data can be collected to determine whether the user's current behavior data deviates from past behavior data. If a deviation occurs, the authentication failure is determined.

[0101] User behavior data can include login records, usage data, etc. Feature vectors corresponding to this user behavior data can be extracted, and these feature vectors can represent user profiles. Then, these feature vectors from past user behavior data after authentication are stored in a database. When a user authenticates, their behavior data from the period preceding authentication is retrieved, the corresponding feature vector is extracted, and the similarity between this feature vector and existing feature vectors in the database is used to determine whether the user's current behavior deviates from past behavior.

[0102] This specification also provides a device for verifying identity, such as... Figure 3 As shown, it includes:

[0103] The dynamic gesture data acquisition module 310 is used to acquire the dynamic gesture data of the user collected during the user's authentication process; the dynamic gesture data is determined based on at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal characteristics of the trajectory of the dynamic gesture.

[0104] The identity verification module 320 is used to verify the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data; the dynamic gesture template is generated based on the user's dynamic gesture data collected during user registration.

[0105] The implementation of the above device is the same as that of the identity verification method described above, and will not be repeated here.

[0106] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must also be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0107] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0108] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions. A typical implementation device is a server system. Of course, this specification does not exclude the possibility that, with the future development of computer technology, the computer implementing the functions of the above embodiments may be, for example, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0109] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive means. The order of steps listed in the embodiments is merely one possible order of execution among many steps and does not represent the only possible order. In actual device or end product execution, the methods shown in the embodiments or drawings may be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment, or even a distributed data processing environment). The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes said elements is not excluded. For example, the use of terms such as "first," "second," etc., is to denote names and does not indicate any particular order.

[0110] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, when implementing one or more of these specifications, the functions of each module can be implemented in one or more software and / or hardware components, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between devices or units, and may be electrical, mechanical, or other forms.

[0111] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0112] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0113] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0114] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0115] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0116] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage, graphene storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0117] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, one or more embodiments of this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0118] One or more embodiments of this specification can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a particular task or implement a particular abstract data type. One or more embodiments of this specification can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0119] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, system embodiments are basically similar to method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. In the description of this specification, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this specification. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described can be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0120] The above description is merely an embodiment of one or more embodiments of this specification and is not intended to limit the scope of these embodiments. Various modifications and variations can be made to these embodiments by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims.

Claims

1. A method for verifying identity, comprising: Acquire dynamic gesture data of the user collected during user authentication; The dynamic gesture data is determined based on at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal characteristics of the trajectory of the dynamic gesture. The user's identity is verified based on the similarity between the stored dynamic gesture template and the dynamic gesture data; the dynamic gesture template is generated based on the user's dynamic gesture data collected during user registration.

2. The method according to claim 1, wherein acquiring the user's dynamic gesture data collected during user authentication includes: The user's dynamic gesture video is captured using a first camera; Based on the dynamic gesture video, obtain the dynamic gesture data; The feature vector corresponding to the dynamic gesture video is obtained by feature extraction of the dynamic gesture video. The trajectory of the dynamic gesture and the temporal features of the dynamic gesture trajectory are obtained by tracking several key points of the hand in the dynamic gesture video. The trajectory of the dynamic gesture is a two-dimensional trajectory.

3. The method according to claim 1, wherein the dynamic gesture data is acquired through the following methods: The system captures dynamic gesture videos of the user using a first camera, as well as 3D dynamic gesture videos captured by a depth camera. The dynamic gesture data is obtained based on the dynamic gesture video and the 3D dynamic gesture video; The feature vector corresponding to the dynamic gesture video is obtained by feature extraction from the dynamic gesture video. The trajectory of the dynamic gesture and the temporal features of the dynamic gesture trajectory are obtained by tracking several key points of the hand in the dynamic gesture video and the three-dimensional dynamic gesture video. The trajectory of the dynamic gesture is a three-dimensional trajectory.

4. The method according to claim 2 or 3, wherein acquiring the dynamic gesture data includes: The extracted first data is used as the dynamic gesture data, and the first data includes at least one of the following: the feature vector corresponding to the dynamic gesture video, the trajectory of the dynamic gesture, and the temporal features of the trajectory of the dynamic gesture. Alternatively, the dynamic gesture data can be obtained by weighted fusion of at least two types of extracted first data.

5. The method according to claim 2 or 3, wherein the method is applied to AI glasses, and further comprises: The user's posture data is collected by the inertial measurement unit of the AI ​​glasses; The dynamic gesture data is compensated based on the posture data to eliminate the influence of the user's head movement from the dynamic gesture data.

6. The method according to claim 2 or 3, further comprising: Show the user the target challenge gestures that need to be performed; The AI ​​glasses use their first camera to capture videos of the user's challenge gestures. If the challenge gesture video matches the target challenge gesture, a continuity detection is performed on the challenge gesture video and the dynamic gesture video to determine whether the challenge gesture video and the dynamic gesture video were filmed at the same location; The step of verifying the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data includes: If the continuity detection determines that the challenge gesture video and the dynamic gesture video were filmed at the same location, the user's identity is verified based on the similarity between the stored dynamic gesture template and the dynamic gesture data.

7. The method according to claim 1, wherein multiple dynamic gesture templates are pre-stored, and each dynamic gesture template is associated with a corresponding static gesture template, wherein the static gesture template is generated based on the static gesture image of the user collected during user registration; Also includes: The static gesture images of the user captured by the first camera are used to obtain the static gesture data of the user collected during the user's authentication process. The step of verifying the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data includes: Determine a target static gesture template that matches the user's static gesture data from the stored static gesture templates, and use the dynamic gesture template associated with the target static gesture template as the target dynamic gesture template; The user's identity is verified based on the similarity between the target dynamic gesture template and the dynamic gesture data.

8. The method according to claim 7, further comprising, in the case that no target static gesture template matches the user's static gesture data, including: Authentication failed.

9. The method according to claim 1, wherein the dynamic gesture template is stored in encrypted form; Before verifying the user's identity based on the similarity between the stored dynamic gesture template and the dynamic gesture data, the process further includes: Decrypt the encrypted dynamic gesture template.

10. A computing device comprising a memory and a processor, wherein the memory stores executable code, and the processor, when executing the executable code, implements the method of any one of claims 1-9.