Posture authentication method, device and equipment, storage medium and computer program product

By extracting and fusing feature distances in video sequences and event stream data, the defects of depth images in existing pose authentication techniques in dynamic information modeling are solved, and the accuracy of pose authentication is improved.

CN119649477BActive Publication Date: 2025-06-06LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510162944.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

In the existing pose authentication technology, deep images have defects in dynamic information modeling, resulting in low accuracy of pose authentication.

Method used

By obtaining the video sequence and event stream data of the standard pose and the pose to be authenticated, input it into the video feature extraction model and the event feature extraction model respectively, extracting video features and event features, calculating feature distances and fusion to improve the accuracy of pose authentication.

Benefits of technology

By fusing multimodal information in video sequences and event stream data, it is possible to more accurately identify and distinguish nuances of postures, improving the accuracy of posture authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119649477B_ABST
    Figure CN119649477B_ABST
Patent Text Reader

Abstract

The present invention discloses a posture authentication method, device and equipment, storage medium and computer program product, which relates to the field of computer technology. The method comprises: obtaining a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated; extracting a first video feature of the first video sequence and a second video feature of the second video sequence; extracting a first event feature of the first event stream data and a second event feature of the second event stream data; calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, fusing the first feature distance and the second feature distance to obtain a fused distance; judging whether the fused distance is greater than a distance threshold; if so, determining that the authentication is passed, and if not, determining that the authentication fails. The present invention improves the accuracy of posture authentication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and more specifically, to a posture authentication method, device and equipment, storage medium and computer program product. Background Art

[0002] Biometric authentication is a technology that uses an individual's unique physiological or behavioral characteristics to verify personal identity. Compared with traditional passwords or tokens, biometrics are unique and difficult to copy. In addition, users do not need to remember complex passwords or carry physical tokens. They can use their own biometrics for identity authentication, which simplifies the authentication process. Authentication using posture, also known as posture biometrics, is a special method in the field of biometrics. It authenticates an individual based on their unique body posture, movement or behavior pattern.

[0003] In related technologies, multimodal methods are mainly used for posture authentication, integrating relevant information of RGB (Red, Green, Blue) images and depth images. However, depth images mainly provide depth information and have defects in modeling dynamic information, resulting in low accuracy of posture authentication.

[0004] Therefore, how to improve the accuracy of posture authentication is a technical problem that those skilled in the art need to solve. Summary of the invention

[0005] The object of the present invention is to provide a posture authentication method, device and equipment, storage medium and computer program product, which improve the accuracy of posture authentication.

[0006] To achieve the above object, the present invention provides a posture authentication method, comprising:

[0007] Acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated;

[0008] Inputting the first video sequence and the second video sequence into the video feature extraction model respectively to extract a first video feature of the first video sequence and a second video feature of the second video sequence;

[0009] Inputting the first event stream data and the second event stream data into the event feature extraction model respectively to extract the first event feature of the first event stream data and the second event feature of the second event stream data;

[0010] Calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and fusing the first feature distance and the second feature distance to obtain a fused distance;

[0011] Determine whether the fusion distance is less than the distance threshold; if so, the authentication is determined to be successful; if not, the authentication is determined to be unsuccessful.

[0012] Among them, it also includes:

[0013] Obtain a first training set; wherein the first training set includes a training video sequence of training posture samples and corresponding labels;

[0014] Training the expanded three-dimensional network based on the first training set to obtain a trained expanded three-dimensional network;

[0015] The classifier in the trained expanded 3D network is removed to obtain the video feature extraction model.

[0016] Among them, the video feature extraction model includes a first part, a second part, a third part, a fourth part and a fifth part connected in sequence. The first part is a three-dimensional convolutional layer whose size is larger than a preset size. The second part includes a maximum pooling layer connected in sequence and two three-dimensional convolutional layers of different sizes. The third part includes a maximum pooling layer connected in sequence and two multi-branch comprehensive structures. The fourth part includes a maximum pooling layer connected in sequence and four multi-branch comprehensive structures. The fifth part includes a maximum pooling layer and a global average pooling layer connected in sequence.

[0017] Among them, the multi-branch comprehensive structure includes a first branch, a second branch, a third branch and a fourth branch distributed in parallel, the first branch includes a three-dimensional convolutional layer, the second branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the third branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the second branch and the third branch contain different numbers of channels of the three-dimensional convolutional layers, and the fourth branch includes a maximum pooling layer and a three-dimensional convolutional layer connected in sequence.

[0018] Among them, it also includes:

[0019] Obtain a second training set; wherein the second training set includes training event stream data of training posture samples and corresponding labels;

[0020] Training a spiking neural network based on a self-attention mechanism based on the second training set to obtain a trained spiking neural network based on a self-attention mechanism;

[0021] The event feature extraction model is obtained by removing the classifier from the trained self-attention mechanism-based spiking neural network.

[0022] Among them, the event feature extraction model includes a tokenization structure connected in sequence and multiple superimposed converter structures.

[0023] Among them, the labeled structure includes a first two-dimensional convolutional layer, a second two-dimensional convolutional layer, a first maximum pooling layer, a third two-dimensional convolutional layer, a second maximum pooling layer, a fourth two-dimensional convolutional layer, a third maximum pooling layer, a fifth two-dimensional convolutional layer, and a fourth maximum pooling layer, which are connected in sequence. The number of channels of the first two-dimensional convolutional layer, the second two-dimensional convolutional layer, the third two-dimensional convolutional layer, the fourth two-dimensional convolutional layer, and the fifth two-dimensional convolutional layer are 16, 32, 64, 128, and 128.

[0024] The converter structure includes a self-attention layer, a first residual connection layer, a multi-layer perceptron layer, and a second residual connection layer connected in sequence;

[0025] The first residual connection layer is used to perform residual connection on the input features and output features of the self-attention layer, and the second residual connection layer is used to perform residual connection on the input features and output features of the multi-layer perceptron layer.

[0026] The self-attention layer includes a conversion structure and a spike conversion structure, the conversion structure includes a one-dimensional convolution layer and a batch normalization layer connected in sequence, and the spike conversion structure includes a one-dimensional convolution layer, a batch normalization layer and a pulse neuron connected in sequence;

[0027] The spike transformation structure is used to construct the feature map of query, key, and value, and the transformation structure is used for feature mapping before the feature output of the self-attention layer.

[0028] The step of calculating a first feature distance between a first video feature and a second video feature, and a second feature distance between a first event feature and a second event feature includes:

[0029] Calculating a first cosine distance between the first video feature and the second video feature, and normalizing the first cosine distance based on a maximum video feature distance and a minimum video feature distance to obtain a first feature distance;

[0030] A second cosine distance between the first event feature and the second event feature is calculated, and the second cosine distance is normalized based on the maximum event feature distance and the minimum event feature distance to obtain a second feature distance.

[0031] Among them, it also includes:

[0032] Obtain a first verification set; wherein the first verification set includes a verification video sequence of a verification posture sample and a corresponding label;

[0033] Constructing a first verification positive sample based on two verification video sequences with the same label, and constructing a first verification negative sample based on two verification video sequences with different labels;

[0034] Inputting the verification video sequences in the first verification positive sample and the first verification negative sample into the video feature extraction model respectively to obtain corresponding verification video features;

[0035] Calculating the video feature distance between the verification video features corresponding to the two verification video sequences in the first verification positive sample, and the video feature distance between the verification video features corresponding to the two verification video sequences in the first verification negative sample;

[0036] Determine the maximum video feature distance and the minimum video feature distance.

[0037] Among them, it also includes:

[0038] Obtain a second verification set; wherein the second verification set includes verification event stream data and corresponding labels of the verification posture samples;

[0039] Construct a second verification positive sample based on two verification event stream data with the same label, and construct a second verification negative sample based on two verification event stream data with different labels;

[0040] Inputting the verification event stream data in the second verification positive sample and the second verification negative sample into the event feature extraction model respectively to obtain corresponding verification event features;

[0041] Calculating the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification positive sample, and the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification negative sample;

[0042] Determine the maximum event feature distance and the minimum event feature distance.

[0043] The fusion distance is obtained by fusing the first feature distance and the second feature distance, including:

[0044] determining a first maximum value of the first characteristic distance and zero, and a second maximum value of the second characteristic distance and zero;

[0045] The sum of the first maximum value and the second maximum value is calculated to obtain the fusion distance.

[0046] Among them, it also includes:

[0047] Obtain a test set; wherein the test set includes test posture samples and corresponding labels;

[0048] Construct a test positive sample based on two test posture samples with the same label, and construct a test negative sample based on two test posture samples with different labels;

[0049] Determine a first test fusion distance between two test pose samples in the test positive sample, and a second test fusion distance between two test pose samples in the test negative sample;

[0050] Determine a plurality of candidate distance thresholds, and determine false rejection rates and false acceptance rates corresponding to the plurality of candidate distance thresholds based on a first test fusion distance of a test positive sample and a second test fusion distance of a test negative sample;

[0051] The candidate distance threshold whose false rejection rate and false acceptance rate are closest is determined as the distance threshold.

[0052] Wherein, determining a first test fusion distance between two test posture samples in a test positive sample and a second test fusion distance between two test posture samples in a test negative sample comprises:

[0053] Input the video sequence of the test posture sample in the test positive sample and the video sequence of the test posture sample in the test negative sample into the video feature extraction model respectively to extract the corresponding test video features;

[0054] The event stream data of the test posture sample in the test positive sample and the event stream data of the test posture sample in the test negative sample are respectively input into the event feature extraction model to extract the corresponding test event features;

[0055] Calculate the test positive sample video feature distance between the test video features of two test posture samples in the test positive sample, and the test positive sample event feature distance between the test event features of two test posture samples in the test positive sample, and fuse the test positive sample video feature distance and the test positive sample event feature distance to obtain a first test fusion distance;

[0056] Calculate the test negative sample video feature distance between the test video features of two test posture samples in the test negative sample, and the test negative sample event feature distance between the test event features of two test posture samples in the test negative sample, and fuse the test negative sample video feature distance and the test negative sample event feature distance to obtain a second test fusion distance.

[0057] The step of obtaining the second video sequence and the second event stream data of the posture to be authenticated includes:

[0058] A second video sequence and second event stream data of the posture to be authenticated are acquired through a dynamic and active pixel visual sensor.

[0059] To achieve the above object, the present invention provides a posture authentication device, comprising:

[0060] An acquisition module, used to acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated;

[0061] A first extraction module, used for inputting the first video sequence and the second video sequence into the video feature extraction model respectively, so as to extract a first video feature of the first video sequence and a second video feature of the second video sequence;

[0062] A second extraction module, used to input the first event stream data and the second event stream data into the event feature extraction model respectively, so as to extract the first event feature of the first event stream data and the second event feature of the second event stream data;

[0063] A fusion module, used to calculate a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and to fuse the first feature distance and the second feature distance to obtain a fusion distance;

[0064] The authentication module is used to determine whether the fusion distance is less than the distance threshold; if so, the authentication is determined to be successful, otherwise, the authentication is determined to be unsuccessful.

[0065] To achieve the above object, the present invention provides an electronic device, comprising:

[0066] Memory for storing computer programs;

[0067] The processor is used to implement the steps of the above-mentioned posture authentication method when executing the computer program.

[0068] To achieve the above object, the present invention provides a non-volatile storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned posture authentication method are implemented.

[0069] To achieve the above object, the present invention provides a computer program product, including a computer program, which implements the steps of the above posture authentication method when executed by a processor.

[0070] It can be seen from the above scheme that a posture authentication method provided by the present invention includes: obtaining a first video sequence and a first event stream data of a standard posture, and a second video sequence and a second event stream data of a posture to be authenticated; inputting the first video sequence and the second video sequence into a video feature extraction model respectively to extract a first video feature of the first video sequence and a second video feature of the second video sequence; inputting the first event stream data and the second event stream data into an event feature extraction model respectively to extract a first event feature of the first event stream data and a second event feature of the second event stream data; calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and fusing the first feature distance and the second feature distance to obtain a fused distance; judging whether the fused distance is greater than a distance threshold; if so, determining that the authentication is passed, and if not, determining that the authentication fails.

[0071] The beneficial effect of the present invention is that the posture authentication method provided by the present invention integrates the visual information provided by the video sequence and the dynamic information provided by the event stream data, thereby realizing posture authentication based on multimodal information. The event stream data can record the brightness change events of each pixel in the scene in the form of a timestamp, and can well model the dynamic process of the posture, including information such as speed and acceleration. By integrating these rich dynamic information into the posture authentication, the authentication system can more accurately identify the subtle differences in posture, thereby improving the accuracy of posture authentication. The present invention also discloses a posture authentication device, an electronic device, a non-volatile storage medium and a computer program product, which can also achieve the above technical effects.

[0072] It is to be understood that the foregoing general description and the following detailed description are exemplary only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. The drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific implementation methods, they are used to explain the present disclosure, but do not constitute a limitation to the present disclosure. In the drawings:

[0074] Figure 1 is a flow chart of a posture authentication method according to an exemplary embodiment;

[0075] Figure 2 is a schematic diagram showing a training process of an expanded three-dimensional network in an authentication system according to an exemplary embodiment;

[0076] Figure 3 is a structural diagram of a video feature extraction model according to an exemplary embodiment;

[0077] Figure 4 is a structural diagram of a multi-branch integrated structure according to an exemplary embodiment;

[0078] Figure 5 is a schematic diagram of a training process of a spike neural network based on a self-attention mechanism in an authentication system according to an exemplary embodiment;

[0079] Figure 6 is a structural diagram of a tokenization structure according to an exemplary embodiment;

[0080] Figure 7 is a structural diagram of a converter structure according to an exemplary embodiment;

[0081] Figure 8 is a structural diagram of a self-attention layer according to an exemplary embodiment;

[0082] Fig. 9 is a schematic diagram of a reasoning process of an authentication system according to an exemplary embodiment;

[0083] Fig.10 A schematic diagram of hardware layout of an authentication system according to an exemplary embodiment;

[0084] Fig.11 It is a schematic diagram showing a process of determining a maximum video feature distance and a minimum video feature distance according to an exemplary embodiment;

[0085] Fig.12 is a schematic diagram showing a process of determining a maximum event feature distance and a minimum event feature distance according to an exemplary embodiment;

[0086] Fig.13 is a schematic diagram showing a process of determining a distance threshold according to an exemplary embodiment;

[0087] Fig.14 is a structural diagram of a posture authentication device according to an exemplary embodiment;

[0088] Fig.15 The figure is a structural diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION

[0089] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention. In addition, in the embodiments of the present invention, "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0090] The embodiment of the invention discloses a posture authentication method, which improves the accuracy of posture authentication.

[0091] See also Figure 1 , a flowchart of a posture authentication method according to an exemplary embodiment is shown, as shown in Figure 1 As shown, including:

[0092] S101: Acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated;

[0093] The first video sequence and the first event stream data of the standard posture are pre-collected and annotated, and are used as a "standard template" for comparison benchmarks. They are derived from a strictly defined posture action, such as a user recording a prescribed action in a specific environment to ensure its accuracy and consistency. The first video sequence contains continuous frame information of the standard posture in the RGB image, which can provide the appearance characteristics of the posture, such as the outline of the human body, the texture of clothing, etc. The first event stream data records the timestamp information of the brightness change of each pixel during the standard posture action, and can capture the dynamic characteristics of the posture, such as the speed and acceleration of the action. The first video sequence and the first event stream data of the standard posture can be pre-collected and stored by a dynamic and active pixel vision sensor (Dynamic Vision and Active Sensing, DAVIS), or can be obtained from other devices, which are not specifically limited here.

[0094] The second video sequence and the second event stream data of the posture to be authenticated are the target data that need to be verified. These data may come from different scenes or different individuals, and their acquisition environment and conditions may be different from those of the standard posture data. The second video sequence provides visual information of the posture to be authenticated, which is used for appearance comparison with the standard posture; the second event stream data records the dynamic information of the posture to be authenticated, which is used to analyze the spatiotemporal characteristics of its movements. By simultaneously acquiring the video sequence and event stream data, the posture features can be fully described from both static and dynamic dimensions. The second video sequence and the second event stream data of the posture to be authenticated can be collected simultaneously by dynamic and active pixel vision sensors.

[0095] S102: Inputting the first video sequence and the second video sequence into a video feature extraction model respectively to extract a first video feature of the first video sequence and a second video feature of the second video sequence;

[0096] Video sequences contain rich visual information, but this information is often high-dimensional and redundant. Directly using it for comparison is inefficient and easily interfered by noise. Therefore, it is necessary to extract key information from video sequences through a video feature extraction model to form a low-dimensional and representative feature vector. Video feature extraction models are usually built based on deep learning technologies, such as convolutional neural networks (CNNs), inflated 3D convolutional networks (I3Ds), etc.

[0097] For the first video sequence, the model will analyze the image content frame by frame and extract features related to the standard posture, such as the position of human joints, the overall outline of the posture, the texture of clothing, etc. These features reflect the visual uniqueness of the standard posture and can provide a benchmark for subsequent comparisons. Similarly, the second video sequence will be input into the same video feature extraction model to extract the second video features. This process will ignore the noise information in the video sequence that is not related to the posture, such as background changes, lighting interference, etc., so as to extract feature vectors that can accurately describe the posture to be authenticated. In this way, the video feature extraction model can convert complex video sequence data into concise and discriminative feature representations, providing reliable input for subsequent feature comparisons.

[0098] Taking the inflated three-dimensional network (I3D) as an example, the construction process of the video feature extraction model includes: obtaining a first training set; wherein the first training set includes a training video sequence of training posture samples and corresponding labels; training the inflated three-dimensional network based on the first training set to obtain a trained inflated three-dimensional network; removing the classifier in the trained inflated three-dimensional network to obtain a video feature extraction model.

[0099] In the specific implementation, the first training set is first obtained, which contains the video sequences of training posture samples and their corresponding labels. The video sequences of these training samples cover a variety of posture actions, and the labels are used to indicate the specific posture category corresponding to each video sequence, that is, the posture of which user, providing supervision information for model training. Subsequently, the expanded three-dimensional network is trained based on the first training set. The I3D network is a deep learning architecture specifically used for video analysis. By applying convolution operations in three-dimensional space, it can effectively capture the spatiotemporal features in video sequences. During the training process, the network learns the mapping relationship between the video sequences and labels of the training samples. The Adam (Adaptive Moment Estimation) optimizer can be used to gradually update the network parameters. At the same time, the parameters can be fixed to 100 epochs (training cycles) for training to improve the recognition of posture features. When the training is completed, an expanded three-dimensional network model with classification capabilities is obtained, which can output the corresponding posture category according to the input video sequence. The training process of the expanded three-dimensional network in the authentication system is as follows: Figure 2 shown.

[0100] However, the goal of the video feature extraction model is to extract features rather than directly classify. Therefore, after training is completed, the classifier part of the expanded three-dimensional network needs to be removed. The classifier is usually located at the end of the network and is used to map the extracted features to specific category labels. After removing the classifier, the retained network structure can focus on extracting representative feature vectors from the input video sequence. These feature vectors contain key information in the video sequence, such as the appearance of the posture, the spatiotemporal characteristics of the action, etc., which can be used for subsequent tasks such as posture authentication and action recognition. The video feature extraction model constructed in this way can not only make full use of the powerful feature extraction capability of the expanded three-dimensional network, but also can be flexibly applied to a variety of downstream tasks, providing efficient and effective feature representation for posture analysis.

[0101] As a feasible implementation mode, the video feature extraction model includes a first part, a second part, a third part, a fourth part and a fifth part connected in sequence, the first part is a three-dimensional convolutional layer whose size is larger than a preset size, the second part includes a maximum pooling layer connected in sequence and two three-dimensional convolutional layers of different sizes, the third part includes a maximum pooling layer connected in sequence and two multi-branch comprehensive structures, the fourth part includes a maximum pooling layer connected in sequence and four multi-branch comprehensive structures, and the fifth part includes a maximum pooling layer and a global average pooling layer connected in sequence.

[0102] In the specific implementation, the structure diagram of the video feature extraction model is as follows Figure 3 As shown, overall, it can be divided into five parts. The first part uses a large convolution kernel as input to quickly reduce the size of the feature map, and its size can be 7×7; the second part captures features at different levels through convolution of the maximum pooling layer (MaxPool3d) and two three-dimensional convolution layers (Conv3d) of different scales. The size of the maximum pooling layer can be 3×3, and the sizes of the two three-dimensional convolution layers can be 1×1 and 3×3; the third part uses the maximum pooling layer and two multi-branch comprehensive structures (Mixed) for feature extraction; the fourth part further uses a deeper network structure to enhance the modeling capability of the features, specifically including a maximum pooling layer of size 3×3 and four multi-branch comprehensive structures; the fifth part obtains the output expression features through the maximum pooling layer and the global average pooling layer (GAP), and the size of the maximum pooling layer can be 2×2. It should be noted that Figure 3 The 3D convolutional layer (Conv3d) in includes the normalization operation.

[0103] As a feasible implementation, the multi-branch comprehensive structure includes a first branch, a second branch, a third branch and a fourth branch distributed in parallel, the first branch includes a three-dimensional convolutional layer, the second branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the third branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the second branch and the third branch contain different numbers of channels of the three-dimensional convolutional layers, and the fourth branch includes a maximum pooling layer and a three-dimensional convolutional layer connected in sequence.

[0104] In a specific implementation, the multi-branch synthesis structure is as follows Figure 4 As shown, for the input feature map, four branches are used. The first branch may include a 3D convolution layer of size 1×1, the second branch may include a 3D convolution layer of size 1×1 and a 3D convolution layer of size 3×3, the third branch may include a 3D convolution layer of size 1×1 and a 3D convolution layer of size 3×3, and the fourth branch may include a maximum pooling layer of size 3×3 and a 3D convolution layer of size 1×1. Among them, the number of channels used by the second and third branches is different. This setting is also for extracting features from different angles. The single convolution of the first branch and the maximum pooling and convolution of the fourth branch are combined. The final output is spliced ​​together in the channel dimension and sent to the next module.

[0105] S103: inputting the first event stream data and the second event stream data into an event feature extraction model respectively to extract a first event feature of the first event stream data and a second event feature of the second event stream data;

[0106] Event stream data records the brightness change events of each pixel in the scene in the form of timestamps. Event stream data focuses on capturing dynamic information and can provide spatiotemporal characteristics such as speed and acceleration of gestures. The function of the event feature extraction model is to extract features related to the dynamic characteristics of gestures from these complex event data. Event feature extraction models are usually built based on deep learning technologies, such as convolutional neural networks (CNNs), spiking neural networks based on self-attention mechanisms, etc.

[0107] For the first event stream data, the event feature extraction model will analyze the time distribution and spatial distribution of the event, and extract the dynamic features of the standard posture action, such as the rhythm of the action, the duration of the key action, etc. These features can accurately describe the dynamic characteristics of the standard posture, and provide a dynamic benchmark for subsequent comparisons. Similarly, the second event stream data will also be input into the event feature extraction model to extract the second event features. This process will focus on the dynamic changes in the event stream, ignoring the background event noise that is not related to the posture action, so as to extract the feature vector that can accurately describe the dynamic characteristics of the posture to be authenticated. Through event feature extraction, the dynamic information in the event stream data can be converted into a concise and discriminative feature representation, further enriching the feature dimension of posture authentication.

[0108] Taking the spiking neural network based on the self-attention mechanism as an example, the construction process of the event feature extraction model includes: obtaining a second training set; wherein the second training set includes training event stream data and corresponding labels of training posture samples; training the spiking neural network based on the self-attention mechanism based on the second training set to obtain a trained spiking neural network based on the self-attention mechanism; removing the classifier in the trained spiking neural network based on the self-attention mechanism to obtain the event feature extraction model.

[0109] In the specific implementation, the second training set is first obtained, which contains the event stream data of the training posture samples and their corresponding labels. The event stream data records the timestamp information of the brightness changes of the pixels in the scene, which can reflect the dynamic characteristics of the posture movements, while the labels are used to indicate the specific posture category corresponding to each event stream data, providing a supervision signal for model training. Subsequently, the spiking neural network based on the self-attention mechanism is trained based on the second training set. The spiking neural network is a neural network architecture that simulates the transmission of biological neuron pulse signals. It can effectively process time series data, and is particularly suitable for processing rich spatiotemporal information in event stream data. The introduction of the self-attention mechanism further enhances the model's ability to focus on key dynamic information, enabling the network to automatically focus on the most representative parts of the event stream, such as the starting point, end point or key turning point of the posture movement. Through training, the network learns how to map event stream data to the corresponding posture category, thereby having the ability to recognize the dynamic characteristics of the posture. The training process of the spiking neural network based on the self-attention mechanism in the authentication system is as follows Figure 5 shown.

[0110] After the training is completed, the classifier part of the spike neural network is removed. The classifier is usually located at the end of the network and is used to map the extracted feature vector to a specific posture category. After removing the classifier, the retained network structure can focus on extracting representative feature vectors from the input event stream data. These feature vectors not only contain the dynamic information of the posture action, such as the speed and acceleration of the action, but also highlight the most valuable information for posture authentication or analysis tasks through the optimization of the self-attention mechanism. The final event feature extraction model can efficiently extract key features from the event stream data, providing strong support for subsequent posture analysis tasks.

[0111] As a feasible implementation method, the event feature extraction model includes a tokenization structure connected in sequence and multiple superimposed transformer structures. In the specific implementation, the Spikingformer architecture integrates the Transformer structure and the pulse neural network, making the network more suitable for processing event data while having lower energy consumption. In visual data processing, ViT is a classic structure, which mainly consists of two parts. One part is to tokenize the input image, and the next part is to use Transformer for structural stacking.

[0112] As a feasible implementation method, the labeling structure includes a first two-dimensional convolution layer, a second two-dimensional convolution layer, a first maximum pooling layer, a third two-dimensional convolution layer, a second maximum pooling layer, a fourth two-dimensional convolution layer, a third maximum pooling layer, a fifth two-dimensional convolution layer, and a fourth maximum pooling layer, which are connected in sequence. The number of channels of the first two-dimensional convolution layer, the second two-dimensional convolution layer, the third two-dimensional convolution layer, the fourth two-dimensional convolution layer, and the fifth two-dimensional convolution layer are 16, 32, 64, 128, and 128.

[0113] In a specific implementation, the tokenization structure is as follows Figure 6 As shown in the figure, compared with the original ViT (Vision Transformer) which uses a large-step convolution operation to achieve tokenization, this embodiment uses a maximum pooling layer (Maxpool2d) to obtain 14×14 patches (local areas). The expression of each patch is achieved by progressively increasing the number of channels, so the number of channels of the two-dimensional convolution layer changes in sequence: 16-32-64-128-128, and the size can be 3×3. In this way, the original event data containing only two channels at the two poles can be encoded into a feature expression patch with 128 channels.

[0114] As a feasible implementation, the converter structure includes a self-attention layer, a first residual connection layer, a multi-layer perceptron layer, and a second residual connection layer connected in sequence; the first residual connection layer is used to perform residual connections on the input features and output features of the self-attention layer, and the second residual connection layer is used to perform residual connections on the input features and output features of the multi-layer perceptron layer.

[0115] In the specific implementation, the tokenized data needs to go through a series of superimposed Transformer structures. The transformer structure is as follows: Figure 7 As shown in the figure, it includes a self-attention layer and a multilayer perceptron layer, and residual connections are used at the same time. First, the input vector enters the self-attention layer. The self-attention layer enables the model to focus on important features at different positions in the sequence when processing sequence data, thereby capturing the dependencies within the sequence. The self-attention layer generates attention weights by calculating the correlation between the elements in the input vector, and then weighted sums the input vector according to these weights to obtain a new representation. This mechanism allows the model to consider the information of other elements in the sequence when processing each element, thereby enhancing the model's ability to understand sequence data. The output of the self-attention layer is added to the original input vector through a residual connection. Residual connections help alleviate the gradient vanishing problem in deep network training, allowing the network to learn more effectively. The result of the addition is used as the input of the MLP (Multilayer Perceptron) layer. The MLP layer usually contains at least one hidden layer, which can perform nonlinear transformations on the data to further extract features. The output of the MLP layer is also added to the output of the self-attention layer (i.e., the input of the MLP layer) through a residual connection. This design once again takes advantage of the residual connection and helps with network training and optimization. Finally, the vector after two residual connections and nonlinear transformation is used as the output of the module. This output vector has the same dimension as the input vector and can be used for subsequent distance measurement.

[0116] As a feasible implementation method, the self-attention layer includes a conversion structure and a spike conversion structure. The conversion structure includes a one-dimensional convolution layer and a batch normalization layer connected in sequence. The spike conversion structure includes a one-dimensional convolution layer, a batch normalization layer and a pulse neuron connected in sequence. The spike conversion structure is used to construct a feature map of queries, keys and values, and the conversion structure is used to perform feature mapping before the feature output of the self-attention layer.

[0117] In the specific implementation, the self-attention layer is shown in Figure 8. The detailed structure of the self-attention layer combines the traditional transformation structure (Trans) and the spike transformation structure (SpikingTrans) to process the input data and generate the query (Query, Q), key (Key, K) and value (Value, V) mapping required by the self-attention mechanism. Figure 8 In the example, the input data (100 is the time step, B is the number of samples processed per iteration BatchSize, 128 is the number of channels, and 14×14 is the number of patches) first passes through a series of spike conversion structures, each of which contains a one-dimensional convolution layer (Conv1d), a batch normalization layer (BatchNorm), and a spiking neuron (Spiking Neuron) to construct the Q, K, and V mappings. This spiking calculation method simulates the behavior of biological neurons and can process information in a way that is closer to the biological nervous system. In this way, the self-attention layer can maintain spiking when calculating attention, which helps to improve computational efficiency and reduce energy consumption. After generating the Q, K, and V mappings, these mappings are used to calculate the core part of the self-attention mechanism, that is, the attention weights are calculated by the dot product of Q and K, and then these weights are used to weighted sum the V mapping to obtain the output of the self-attention layer. Before the output of the self-attention layer, the features are further mapped using a conversion structure, which also contains a one-dimensional convolution layer and a batch normalization layer, but does not contain spiking neurons. This design allows for finer adjustment and optimization of features while maintaining the advantages of spiking computation. Ultimately, the output of the self-attention layer further enhances the feature expression capability through a series of transformation operations, providing richer information for subsequent processing steps. The entire structure is designed to achieve a more efficient and biological neural system-like information processing method by combining spiking neural networks and traditional deep learning techniques.

[0118] S104: calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and fusing the first feature distance and the second feature distance to obtain a fused distance;

[0119] The first feature distance is used to measure the difference between the first video feature and the second video feature, and can be measured using distance measurement methods such as Euclidean distance and cosine distance. By calculating the first feature distance, the similarity between the posture to be authenticated and the standard posture in visual appearance can be quantified. For example, if the visual features such as the contours and joint positions of the two postures are very similar, then the first feature distance will be small; conversely, if the difference is large, the first feature distance will be large.

[0120] Similarly, the second feature distance is used to measure the difference between the first event feature and the second event feature, which reflects the similarity between the dynamic characteristics of the posture to be authenticated and the standard posture. By calculating the second feature distance, the matching degree of the two postures in the dynamic characteristics such as the speed and acceleration of the movement can be judged. For example, if the movement rhythm of the two postures is consistent, the second feature distance will be smaller; conversely, if the dynamic characteristics are greatly different, the second feature distance will be larger.

[0121] In order to comprehensively consider the comparison results of video features and event features, it is necessary to fuse the first feature distance and the second feature distance to obtain the fused distance. The fusion process can be achieved by weighted summation, taking the maximum or minimum value, etc. The specific method depends on the evaluation of the importance of different features. The fused distance can comprehensively reflect the overall similarity between the posture to be authenticated and the standard posture in terms of appearance and dynamics, providing a comprehensive basis for the final authentication decision.

[0122] As a feasible implementation method, the first feature distance between the first video feature and the second video feature, and the second feature distance between the first event feature and the second event feature are calculated, including: calculating the first cosine distance between the first video feature and the second video feature, and normalizing the first cosine distance based on the maximum video feature distance and the minimum video feature distance to obtain the first feature distance; calculating the second cosine distance between the first event feature and the second event feature, and normalizing the second cosine distance based on the maximum event feature distance and the minimum event feature distance to obtain the second feature distance.

[0123] In a specific implementation, the calculation formula of the first cosine distance is:

[0124] ;in, is the first cosine distance, is the first video feature, is the second video feature.

[0125] The calculation formula for the second cosine distance is:

[0126] ;in, is the second cosine distance, is the first event feature, It is the characteristic of the second event.

[0127] The normalized formula for the first cosine distance is:

[0128] ;in, is the first characteristic distance, is the maximum video feature distance, is the minimum video feature distance.

[0129] The normalized formula for the second cosine distance is:

[0130] ;in, is the second characteristic distance, is the maximum event feature distance, is the minimum event feature distance.

[0131] , , , It can be calculated on the validation set.

[0132] As a feasible implementation, the first feature distance and the second feature distance are fused to obtain the fused distance, including: determining a first maximum value between the first feature distance and zero, and a second maximum value between the second feature distance and zero; and calculating the sum of the first maximum value and the second maximum value to obtain the fused distance.

[0133] In the specific implementation, in order to avoid negative distance values, the following method can be used to fuse the two normalized distances:

[0134] ;

[0135] Among them, d is the fusion distance.

[0136] S105: Determine whether the fusion distance is less than the distance threshold; if so, determine that the authentication is successful; if not, determine that the authentication is unsuccessful.

[0137] The distance threshold is pre-set according to the actual application scenario and authentication requirements. It defines the maximum acceptable difference range between the posture to be authenticated and the standard posture. The EER (Equal Error Rate) of the fusion distance on the test set can be used as an indicator to set the distance threshold. If the fusion distance is less than the distance threshold, it means that the difference between the posture to be authenticated and the standard posture in appearance and dynamic characteristics is small. It can be considered that the posture to be authenticated is highly similar to the standard posture, so the authentication is judged to be passed. For example, in a security authentication scenario, if the fusion distance between the posture to be authenticated and the standard posture is small, it means that the posture meets the preset standard action and can be allowed to pass the authentication. On the contrary, if the fusion distance is greater than or equal to the distance threshold, it means that there is a large difference between the posture to be authenticated and the standard posture, which may be caused by irregular movements, wrong postures, or interference during data collection. In this case, the authentication is judged to have failed to ensure the accuracy and security of posture authentication. By setting the distance threshold and making judgments based on the fusion distance, it is possible to achieve automated decision-making on posture authentication and improve authentication efficiency and accuracy. The reasoning process of the authentication system is as follows: Fig. 9 shown.

[0138] The hardware layout of the authentication system is as follows Fig.10As shown in the figure, firstly, the data is preprocessed on the host computer, including marking different samples, because each sample contains event stream data and RGB video sequence data, and different data need to be marked to accelerate the later data distribution. The host computer transmits the marked data to the FPGA (Field-Programmable Gate Array) board through the bus, and first distributes the data. Due to the limitation of on-chip resources, we send the data that can be processed directly to the video feature extraction model and the event feature extraction model for processing. The modules that cannot be processed in time are temporarily stored in HBM (High Bandwidth Memory). After the two types of data of the corresponding samples are processed, the distance comparison is performed in the distance comparison module and the result is returned to the host computer. After the host computer obtains the distance data, it determines whether the authentication system has passed according to the defined threshold.

[0139] The posture authentication method provided by the embodiment of the present invention integrates the visual information provided by the video sequence and the dynamic information provided by the event stream data, and realizes posture authentication based on multimodal information. The event stream data can record the brightness change event of each pixel in the scene in the form of a timestamp, and can well model the dynamic process of the posture, including information such as speed and acceleration. By integrating this rich dynamic information into the posture authentication, the authentication system can more accurately identify the subtle differences in posture, thereby improving the accuracy of posture authentication.

[0140] This embodiment introduces the process of determining the maximum video feature distance and the minimum video feature distance. Fig.11 As shown, the following steps are included:

[0141] S201: Obtain a first verification set; wherein the first verification set includes a verification video sequence of a verification posture sample and a corresponding label;

[0142] In this step, a data set for verifying the performance of the video feature extraction model is collected and prepared, that is, a first verification set, including verification video sequences of multiple verification posture samples and corresponding labels.

[0143] S202: constructing a first verification positive sample based on two verification video sequences with the same label, and constructing a first verification negative sample based on two verification video sequences with different labels;

[0144] In this step, positive samples and negative samples are constructed. The positive sample consists of two verification video sequences with the same label, which represent the same posture and action, and are used to verify the accuracy of the model in identifying the same action. The negative sample consists of two verification video sequences with different labels, which represent different postures and actions, and are used to test the ability of the model to distinguish different actions. When constructing the first verification positive sample, for each verification video sequence in the first verification set, a verification video sequence with the same label as that in the first verification set can be randomly selected to form the first verification positive sample, that is, the number of verification posture samples contained in the first verification set is equal to the number of pairs of first verification positive samples constructed. Similarly, when constructing the first verification negative sample, for each verification video sequence in the first verification set, a verification video sequence with a different label as that in the first verification set can be randomly selected to form the first verification negative sample, that is, the number of verification posture samples contained in the first verification set is equal to the number of pairs of first verification negative samples constructed.

[0145] S203: inputting the verification video sequences in the first verification positive sample and the first verification negative sample into the video feature extraction model respectively to obtain corresponding verification video features;

[0146] In this step, all video sequences in the constructed positive samples and negative samples are input into the previously trained video feature extraction model to extract the corresponding verification video features.

[0147] S204: Calculating a video feature distance between verification video features corresponding to two verification video sequences in the first verification positive sample, and a video feature distance between verification video features corresponding to two verification video sequences in the first verification negative sample;

[0148] In this step, the distance between video features in positive and negative samples is calculated. For positive samples, ideally, the feature distance of video sequences from the same label should be small, indicating that they are close to each other in the feature space. For negative samples, the feature distance of video sequences from different labels should be large, indicating that they are distinguished from each other in the feature space.

[0149] S205: Determine a maximum video feature distance and a minimum video feature distance.

[0150] In this step, the maximum video feature distance and the minimum video feature distance are determined from all the calculated video feature distances. The maximum video feature distance usually comes from the negative sample, indicating the maximum distinction between different gestures and actions; the minimum video feature distance comes from the positive sample, indicating the minimum similarity between the same gestures and actions.

[0151] This embodiment introduces the process of determining the maximum event feature distance and the minimum event feature distance, such as Fig.12 As shown, the following steps are included:

[0152] S301: Obtain a second verification set; wherein the second verification set includes verification event stream data of verification posture samples and corresponding labels;

[0153] In this step, a data set for verifying the performance of the video feature extraction model is collected and prepared, namely, a second verification set, including verification event stream data of multiple verification posture samples and corresponding labels.

[0154] S302: constructing a second verification positive sample based on two verification event stream data with the same label, and constructing a second verification negative sample based on two verification event stream data with different labels;

[0155] In this step, positive samples and negative samples are constructed. The positive sample is composed of two verification event stream data with the same label, which represent the same posture and action, and are used to verify the accuracy of the model in identifying the same action. The negative sample is composed of two verification event stream data with different labels, which represent different postures and actions, and are used to test the ability of the model to distinguish different actions. When constructing the second verification positive sample, for each verification event stream data in the second verification set, a verification event stream data with the same label can be randomly selected in the second verification set to form the second verification positive sample, that is, the second verification set contains as many verification posture samples as the number of pairs of second verification positive samples. Similarly, when constructing the second verification negative sample, for each verification event stream data in the second verification set, a verification event stream data with a different label can be randomly selected in the second verification set to form the second verification negative sample, that is, the second verification set contains as many verification posture samples as the number of pairs of second verification negative samples.

[0156] S303: inputting the verification event stream data in the second verification positive sample and the second verification negative sample into the event feature extraction model respectively to obtain corresponding verification event features;

[0157] In this step, all event stream data in the constructed positive samples and negative samples are input into the previously trained event feature extraction model to extract the corresponding verification event features.

[0158] S304: Calculate the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification positive sample, and the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification negative sample;

[0159] In this step, the distance between event features in positive and negative samples is calculated. For positive samples, ideally, the feature distance of event stream data from the same label should be small, indicating that they are close to each other in the feature space. For negative samples, the feature distance of event stream data from different labels should be large, indicating that they are distinguished from each other in the feature space.

[0160] S305: Determine a maximum event feature distance and a minimum event feature distance.

[0161] In this step, the maximum event feature distance and the minimum event feature distance are determined from all the calculated event feature distances. The maximum event feature distance usually comes from the negative sample, indicating the maximum distinction between different gestures and actions; the minimum event feature distance comes from the positive sample, indicating the minimum similarity between the same gestures and actions.

[0162] This embodiment introduces the process of determining the distance threshold. Fig.13 As shown, the following steps are included:

[0163] S401: Obtain a test set; wherein the test set includes test posture samples and corresponding labels;

[0164] In the specific implementation, a dataset for testing the performance of the model is collected and prepared. The test set consists of multiple test posture samples and their corresponding labels. These samples cover various posture actions that the model needs to recognize, and the labels provide the correct classification information for each sample. The test set is used to evaluate the generalization ability of the model in practical applications, that is, the model's ability to recognize and classify unseen data.

[0165] S402: constructing a test positive sample based on two test posture samples with the same label, and constructing a test negative sample based on two test posture samples with different labels;

[0166] In this step, positive samples and negative samples are constructed. The test positive sample consists of two test posture samples with the same label, which represent the same posture action and are used to evaluate the accuracy of the model in identifying the same action. The test negative sample consists of two test posture samples with different labels, which represent different posture actions and are used to test the model's ability to distinguish different actions. When testing the positive sample, for each test posture sample in the test set, a test posture sample with the same label as the test posture sample can be randomly selected in the test set to form a test positive sample, that is, the number of test posture samples contained in the test set is the number of pairs of test positive samples constructed. Similarly, when testing the negative sample, for each test posture sample in the test set, a test posture sample with a different label can be randomly selected in the test set to form a test negative sample, that is, the number of test posture samples contained in the test set is the number of pairs of test negative samples constructed.

[0167] S403: Determine a first test fusion distance between two test posture samples in the test positive sample, and a second test fusion distance between two test posture samples in the test negative sample;

[0168] In this step, the fusion distance between the pose samples in the test positive sample and the test negative sample is calculated. The fusion distance is calculated by comprehensively considering the distance of video features and event features, which can more comprehensively reflect the similarity between two pose samples. For the test positive sample, the expected fusion distance is small, indicating that they are close to each other in the feature space; for the test negative sample, the expected fusion distance is large, indicating that they are distinguished from each other in the feature space. The calculation of these fusion distances helps to evaluate the distinguishing ability of the model in practical applications.

[0169] As a feasible implementation method, determining a first test fusion distance between two test posture samples in a test positive sample and a second test fusion distance between two test posture samples in a test negative sample includes: inputting the video sequence of the test posture sample in the test positive sample and the video sequence of the test posture sample in the test negative sample into a video feature extraction model respectively to extract corresponding test video features; inputting the event stream data of the test posture sample in the test positive sample and the event stream data of the test posture sample in the test negative sample into an event feature extraction model respectively to extract corresponding test event features; calculating the test positive sample video feature distance between the test video features of two test posture samples in the test positive sample and the test positive sample event feature distance between the test event features of two test posture samples in the test positive sample, and fusing the test positive sample video feature distance and the test positive sample event feature distance to obtain a first test fusion distance; calculating the test negative sample video feature distance between the test video features of two test posture samples in the test negative sample and the test negative sample event feature distance between the test event features of two test posture samples in the test negative sample, and fusing the test negative sample video feature distance and the test negative sample event feature distance to obtain a second test fusion distance.

[0170] In the specific implementation, first, the video sequence corresponding to each posture sample in the test positive sample and the test negative sample is input into the video feature extraction model to extract the corresponding test video features. Secondly, the event stream data corresponding to each posture sample in the test positive sample and the test negative sample is input into the event feature extraction model to extract the test event features. For the test positive sample, the distance between the test video features of the two posture samples (test positive sample video feature distance) and the distance between the test event features (test positive sample event feature distance) are calculated respectively. Then, these two distances are fused to obtain the first test fusion distance. This fusion distance reflects the similarity between the two posture samples with the same label in the feature space. For the test negative sample, the distance between the test video features of the two posture samples (test negative sample video feature distance) and the distance between the test event features (test negative sample event feature distance) are also calculated respectively. Then, these two distances are fused to obtain the second test fusion distance. This fusion distance reflects the dissimilarity between the two posture samples with different labels in the feature space.

[0171] S404: Determine multiple candidate distance thresholds, and determine false rejection rates and false acceptance rates corresponding to the multiple candidate distance thresholds based on a first test fusion distance of a test positive sample and a second test fusion distance of a test negative sample;

[0172] In this step, multiple candidate distance thresholds are first determined. For each candidate distance threshold, the corresponding false rejection rate (FRR) and false acceptance rate (FAR) are calculated respectively. The false rejection rate indicates the proportion of positive samples that are mistakenly judged as negative samples, while the false acceptance rate indicates the proportion of negative samples that are mistakenly judged as positive samples. The specific calculation process is as follows: the first test fusion distance of the test positive sample is compared with the candidate distance threshold. If the first test fusion distance is less than the candidate distance threshold, the judgment is correct. If the first test fusion distance is greater than or equal to the candidate distance threshold, the judgment is wrong. The proportion of the wrongly judged positive samples to the total number of positive samples is counted as the false rejection rate. The second test fusion distance of the test negative sample is compared with the candidate distance threshold. If the second test fusion distance is less than the candidate distance threshold, the judgment is wrong. If the second test fusion distance is greater than or equal to the candidate distance threshold, the judgment is correct. The proportion of the wrongly judged negative samples to the total number of negative samples is counted as the false acceptance rate. By analyzing the false rejection rate and false acceptance rate corresponding to different candidate thresholds, the performance of the model under different threshold settings can be evaluated.

[0173] S405: Determine the candidate distance threshold whose false rejection rate and false acceptance rate are closest as the distance threshold.

[0174] In this step, an optimal distance threshold is selected from multiple candidate distance thresholds, and the false rejection rate and false acceptance rate corresponding to the distance threshold are closest. This optimal distance threshold can reduce the false rejection rate as much as possible while ensuring a low false acceptance rate, thereby achieving the best balance in practical applications. By selecting this optimal distance threshold, the accuracy and reliability of the model can be improved, ensuring that it can correctly identify and classify gestures in practical applications.

[0175] A posture authentication device provided by an embodiment of the present invention is introduced below. The posture authentication device described below and the posture authentication method described above can be referenced to each other.

[0176] See also Fig.14 , a structural diagram of a posture authentication device according to an exemplary embodiment is shown, such as Fig.14 As shown, including:

[0177] An acquisition module 100 is used to acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated;

[0178] A first extraction module 200, for inputting the first video sequence and the second video sequence into a video feature extraction model respectively, to extract a first video feature of the first video sequence and a second video feature of the second video sequence;

[0179] A second extraction module 300, for inputting the first event stream data and the second event stream data into an event feature extraction model respectively, so as to extract a first event feature of the first event stream data and a second event feature of the second event stream data;

[0180] A fusion module 400 is used to calculate a first feature distance between a first video feature and a second video feature, and a second feature distance between a first event feature and a second event feature, and fuse the first feature distance and the second feature distance to obtain a fusion distance;

[0181] The authentication module 500 is used to determine whether the fusion distance is less than a distance threshold; if so, the authentication is determined to be successful; if not, the authentication is determined to be unsuccessful.

[0182] The posture authentication device provided by the embodiment of the present invention integrates the visual information provided by the video sequence and the dynamic information provided by the event stream data, and realizes posture authentication based on multimodal information. The event stream data can record the brightness change event of each pixel in the scene in the form of a timestamp, and can well model the dynamic process of the posture, including information such as speed and acceleration. By integrating this rich dynamic information into the posture authentication, the authentication system can more accurately identify the subtle differences in posture, thereby improving the accuracy of posture authentication.

[0183] Based on the above embodiment, as a preferred implementation, it also includes:

[0184] The first training module is used to obtain a first training set; wherein the first training set includes a training video sequence of training posture samples and corresponding labels; based on the first training set, an expanded three-dimensional network is trained to obtain a trained expanded three-dimensional network; and a classifier in the trained expanded three-dimensional network is removed to obtain a video feature extraction model.

[0185] On the basis of the above embodiments, as a preferred implementation mode, the video feature extraction model includes a first part, a second part, a third part, a fourth part and a fifth part connected in sequence, the first part is a three-dimensional convolutional layer whose size is greater than a preset size, the second part includes a maximum pooling layer connected in sequence and two three-dimensional convolutional layers of different sizes, the third part includes a maximum pooling layer connected in sequence and two multi-branch comprehensive structures, the fourth part includes a maximum pooling layer connected in sequence and four multi-branch comprehensive structures, and the fifth part includes a maximum pooling layer and a global average pooling layer connected in sequence.

[0186] Based on the above embodiments, as a preferred implementation, the multi-branch integrated structure includes a first branch, a second branch, a third branch and a fourth branch distributed in parallel, the first branch includes a three-dimensional convolutional layer, the second branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the third branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the second branch and the third branch contain different numbers of channels of the three-dimensional convolutional layers, and the fourth branch includes a maximum pooling layer and a three-dimensional convolutional layer connected in sequence.

[0187] Based on the above embodiment, as a preferred implementation, it also includes:

[0188] The second training module is used to obtain a second training set; wherein the second training set includes training event stream data of training posture samples and corresponding labels; based on the second training set, a spiking neural network based on a self-attention mechanism is trained to obtain a trained spiking neural network based on a self-attention mechanism; and the classifier in the trained spiking neural network based on a self-attention mechanism is removed to obtain an event feature extraction model.

[0189] Based on the above embodiments, as a preferred implementation, the event feature extraction model includes a tokenization structure connected in sequence and a plurality of superimposed converter structures.

[0190] Based on the above embodiments, as a preferred implementation, the labeling structure includes a first two-dimensional convolutional layer, a second two-dimensional convolutional layer, a first maximum pooling layer, a third two-dimensional convolutional layer, a second maximum pooling layer, a fourth two-dimensional convolutional layer, a third maximum pooling layer, a fifth two-dimensional convolutional layer, and a fourth maximum pooling layer, which are connected in sequence, and the number of channels of the first two-dimensional convolutional layer, the second two-dimensional convolutional layer, the third two-dimensional convolutional layer, the fourth two-dimensional convolutional layer, and the fifth two-dimensional convolutional layer are distributed as 16, 32, 64, 128, and 128.

[0191] Based on the above embodiments, as a preferred implementation, the converter structure includes a self-attention layer, a first residual connection layer, a multi-layer perceptron layer, and a second residual connection layer connected in sequence; the first residual connection layer is used to perform residual connections on the input features and output features of the self-attention layer, and the second residual connection layer is used to perform residual connections on the input features and output features of the multi-layer perceptron layer.

[0192] Based on the above embodiments, as a preferred implementation, the self-attention layer includes a conversion structure and a spike conversion structure, the conversion structure includes a one-dimensional convolutional layer and a batch normalization layer connected in sequence, and the spike conversion structure includes a one-dimensional convolutional layer, a batch normalization layer and a pulse neuron connected in sequence; the spike conversion structure is used to construct a feature map of queries, keys, and values, and the conversion structure is used to perform feature mapping before the feature output of the self-attention layer.

[0193] On the basis of the above embodiments, as a preferred implementation mode, the fusion module 400 is specifically used to: calculate the first cosine distance between the first video feature and the second video feature, and normalize the first cosine distance based on the maximum video feature distance and the minimum video feature distance to obtain the first feature distance; calculate the second cosine distance between the first event feature and the second event feature, and normalize the second cosine distance based on the maximum event feature distance and the minimum event feature distance to obtain the second feature distance.

[0194] Based on the above embodiment, as a preferred implementation, it also includes:

[0195] A first verification module is used to obtain a first verification set; wherein the first verification set includes a verification video sequence and corresponding labels of a verification posture sample; a first verification positive sample is constructed based on two verification video sequences with the same label, and a first verification negative sample is constructed based on two verification video sequences with different labels; the verification video sequences in the first verification positive sample and the first verification negative sample are respectively input into a video feature extraction model to obtain corresponding verification video features; the video feature distance between the verification video features corresponding to the two verification video sequences in the first verification positive sample and the video feature distance between the verification video features corresponding to the two verification video sequences in the first verification negative sample are calculated; and the maximum video feature distance and the minimum video feature distance are determined.

[0196] Based on the above embodiment, as a preferred implementation, it also includes:

[0197] The second verification module is used to obtain a second verification set; wherein the second verification set includes verification event stream data and corresponding labels of the verification posture samples; construct a second verification positive sample based on two verification event stream data with the same label, and construct a second verification negative sample based on two verification event stream data with different labels; input the verification event stream data in the second verification positive sample and the second verification negative sample into the event feature extraction model respectively to obtain corresponding verification event features; calculate the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification positive sample, and the event feature distance between the verification event features corresponding to the two verification event stream data in the second verification negative sample; determine the maximum event feature distance and the minimum event feature distance.

[0198] Based on the above embodiment, as a preferred implementation, the fusion module 400 is specifically used to: determine the first maximum value between the first feature distance and zero, and the second maximum value between the second feature distance and zero; calculate the sum of the first maximum value and the second maximum value to obtain the fusion distance.

[0199] Based on the above embodiment, as a preferred implementation, it also includes:

[0200] A test module is used to obtain a test set; wherein the test set includes test posture samples and corresponding labels; construct a test positive sample based on two test posture samples with the same label, and construct a test negative sample based on two test posture samples with different labels; determine a first test fusion distance between two test posture samples in the test positive sample, and a second test fusion distance between two test posture samples in the test negative sample; determine multiple candidate distance thresholds, and determine false rejection rates and false pass rates corresponding to multiple candidate distance thresholds based on the first test fusion distance of the test positive sample and the second test fusion distance of the test negative sample; determine the candidate distance threshold with the closest false rejection rate and false pass rate as the distance threshold.

[0201] On the basis of the above embodiments, as a preferred implementation mode, the test module is specifically used to: input the video sequence of the test posture sample in the test positive sample and the video sequence of the test posture sample in the test negative sample into the video feature extraction model respectively to extract the corresponding test video features; input the event stream data of the test posture sample in the test positive sample and the event stream data of the test posture sample in the test negative sample into the event feature extraction model respectively to extract the corresponding test event features; calculate the test positive sample video feature distance between the test video features of two test posture samples in the test positive sample, and the test positive sample event feature distance between the test event features of two test posture samples in the test positive sample, and fuse the test positive sample video feature distance and the test positive sample event feature distance to obtain a first test fusion distance; calculate the test negative sample video feature distance between the test video features of two test posture samples in the test negative sample, and the test negative sample event feature distance between the test event features of two test posture samples in the test negative sample, and fuse the test negative sample video feature distance and the test negative sample event feature distance to obtain a second test fusion distance.

[0202] Based on the above embodiment, as a preferred implementation, the acquisition module 100 is specifically used to: acquire the second video sequence and the second event stream data of the posture to be authenticated through the dynamic and active pixel vision sensor.

[0203] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0204] Based on the hardware implementation of the above program modules and in order to implement the method of the embodiment of the present invention, the embodiment of the present invention further provides an electronic device, Fig.15 FIG. 1 is a structural diagram of an electronic device according to an exemplary embodiment. Fig.15 As shown, the electronic equipment includes:

[0205] Communication interface 1, capable of exchanging information with other devices such as network devices;

[0206] The processor 2 is connected to the communication interface 1 to realize information exchange with other devices and is used to execute the posture authentication method provided by one or more technical solutions when running a computer program. The computer program is stored in the memory 3.

[0207] Of course, in actual application, the various components in the electronic device are coupled together through the bus system 4. It can be understood that the bus system 4 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 4 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Fig.15 Various buses are labeled as bus system 4 .

[0208] The memory 3 in the embodiment of the present invention is used to store various types of data to support the operation of the electronic device. Examples of such data include: any computer program used to operate on the electronic device.

[0209] It can be understood that the memory 3 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memory 3 described in the embodiments of the present invention is intended to include but is not limited to these and any other suitable types of memories.

[0210] The method disclosed in the above embodiment of the present invention can be applied to the processor 2, or implemented by the processor 2. The processor 2 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 2 or the instruction in the form of software. The above processor 2 can be a general-purpose processor, a DSP, or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 2 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiment of the present invention. The general-purpose processor can be a microprocessor or any conventional processor, etc. In combination with the steps of the method disclosed in the embodiment of the present invention, it can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory 3. The processor 2 reads the program in the memory 3 and completes the steps of the above method in combination with its hardware.

[0211] When the processor 2 executes the program, the corresponding processes in the various methods of the embodiments of the present invention are implemented, which will not be described in detail here for the sake of brevity.

[0212] In an exemplary embodiment, the present invention further provides a non-volatile storage medium storing a computer program, which can be executed by the processor 2 to complete the aforementioned method steps.

[0213] In an exemplary embodiment, the present invention further provides a computer program product, including a computer program, which is executed by the processor 2 to complete the aforementioned method steps.

[0214] A person of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to computer program instructions, and the aforementioned computer program can be stored in a non-volatile storage medium. When the computer program is executed, it executes the steps of the above method embodiments. Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a non-volatile storage medium and includes a number of instructions for an electronic device (which can be a personal computer, a server, a network device, etc.) to execute all or part of the methods of each embodiment of the present invention.

[0215] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A posture authentication method, characterized in that: include: Acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated; Inputting the first video sequence and the second video sequence into a video feature extraction model respectively to extract a first video feature of the first video sequence and a second video feature of the second video sequence; wherein the video feature extraction model is a model obtained by removing the classifier in the trained expanded three-dimensional network; The first event stream data and the second event stream data are respectively input into an event feature extraction model to extract a first event feature of the first event stream data and a second event feature of the second event stream data; wherein the event feature extraction model is a model obtained by removing a classifier in a trained spike neural network based on a self-attention mechanism, and the event feature extraction model includes a sequentially connected tokenization structure and a plurality of superimposed converter structures; Calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and fusing the first feature distance and the second feature distance to obtain a fused distance; Determine whether the fusion distance is less than a distance threshold; if so, determine that the authentication is successful; if not, determine that the authentication fails.

2. The posture authentication method according to claim 1, characterized in that: Also includes: Acquire a first training set; wherein the first training set includes a training video sequence of training posture samples and corresponding labels; Training the expanded three-dimensional network based on the first training set to obtain a trained expanded three-dimensional network; The video feature extraction model is obtained by removing the classifier in the trained expanded three-dimensional network.

3. The posture authentication method according to claim 2, characterized in that: The video feature extraction model includes a first part, a second part, a third part, a fourth part and a fifth part connected in sequence, the first part is a three-dimensional convolutional layer whose size is larger than a preset size, the second part includes a maximum pooling layer connected in sequence and two three-dimensional convolutional layers of different sizes, the third part includes a maximum pooling layer connected in sequence and two multi-branch comprehensive structures, the fourth part includes a maximum pooling layer connected in sequence and four multi-branch comprehensive structures, and the fifth part includes a maximum pooling layer and a global average pooling layer connected in sequence.

4. The posture authentication method according to claim 3, characterized in that: The multi-branch integrated structure includes a first branch, a second branch, a third branch and a fourth branch distributed in parallel, the first branch includes a three-dimensional convolutional layer, the second branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the third branch includes two three-dimensional convolutional layers of different sizes connected in sequence, the second branch and the third branch contain different numbers of channels of the three-dimensional convolutional layers, and the fourth branch includes a maximum pooling layer and a three-dimensional convolutional layer connected in sequence.

5. The posture authentication method according to claim 1, characterized in that: Also includes: Acquire a second training set; wherein the second training set includes training event stream data of training posture samples and corresponding labels; Training a spiking neural network based on a self-attention mechanism based on the second training set to obtain a trained spiking neural network based on a self-attention mechanism; The event feature extraction model is obtained by removing the classifier in the trained self-attention mechanism-based spike neural network.

6. The posture authentication method according to claim 1, characterized in that: The labeling structure includes a first two-dimensional convolutional layer, a second two-dimensional convolutional layer, a first maximum pooling layer, a third two-dimensional convolutional layer, a second maximum pooling layer, a fourth two-dimensional convolutional layer, a third maximum pooling layer, a fifth two-dimensional convolutional layer, and a fourth maximum pooling layer, which are connected in sequence. The number of channels of the first two-dimensional convolutional layer, the second two-dimensional convolutional layer, the third two-dimensional convolutional layer, the fourth two-dimensional convolutional layer, and the fifth two-dimensional convolutional layer are 16, 32, 64, 128, and 128.

7. The posture authentication method according to claim 1, characterized in that: The converter structure includes a self-attention layer, a first residual connection layer, a multi-layer perceptron layer, and a second residual connection layer connected in sequence; The first residual connection layer is used to perform residual connection on the input features and output features of the self-attention layer, and the second residual connection layer is used to perform residual connection on the input features and output features of the multi-layer perceptron layer.

8. The posture authentication method according to claim 7, characterized in that: The self-attention layer includes a conversion structure and a spike conversion structure, the conversion structure includes a one-dimensional convolution layer and a batch normalization layer connected in sequence, and the spike conversion structure includes a one-dimensional convolution layer, a batch normalization layer and a pulse neuron connected in sequence; The spike conversion structure is used to construct a feature map of query, key, and value, and the conversion structure is used to perform feature mapping before the feature output of the self-attention layer.

9. The posture authentication method according to claim 1, characterized in that: Calculating a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, comprising: Calculating a first cosine distance between the first video feature and the second video feature, and normalizing the first cosine distance based on a maximum video feature distance and a minimum video feature distance to obtain a first feature distance; A second cosine distance between the first event feature and the second event feature is calculated, and the second cosine distance is normalized based on a maximum event feature distance and a minimum event feature distance to obtain a second feature distance.

10. The posture authentication method according to claim 9, characterized in that: Also includes: Obtain a first verification set; wherein the first verification set includes a verification video sequence of a verification posture sample and a corresponding label; Constructing a first verification positive sample based on two verification video sequences with the same label, and constructing a first verification negative sample based on two verification video sequences with different labels; Inputting the verification video sequences in the first verification positive sample and the first verification negative sample into the video feature extraction model respectively to obtain corresponding verification video features; Calculating a video feature distance between verification video features corresponding to two verification video sequences in the first verification positive sample, and a video feature distance between verification video features corresponding to two verification video sequences in the first verification negative sample; The maximum video feature distance and the minimum video feature distance are determined.

11. The posture authentication method according to claim 9, characterized in that: Also includes: Obtain a second verification set; wherein the second verification set includes verification event stream data and corresponding labels of verification posture samples; Construct a second verification positive sample based on two verification event stream data with the same label, and construct a second verification negative sample based on two verification event stream data with different labels; Inputting the verification event stream data in the second verification positive sample and the second verification negative sample into the event feature extraction model respectively to obtain corresponding verification event features; Calculating an event feature distance between verification event features corresponding to two verification event stream data in the second verification positive sample, and an event feature distance between verification event features corresponding to two verification event stream data in the second verification negative sample; The maximum event feature distance and the minimum event feature distance are determined.

12. The posture authentication method according to claim 1, characterized in that: Fusion of the first characteristic distance and the second characteristic distance to obtain a fusion distance includes: determining a first maximum value of the first characteristic distance and zero, and a second maximum value of the second characteristic distance and zero; The sum of the first maximum value and the second maximum value is calculated to obtain a fusion distance.

13. The posture authentication method according to claim 1, characterized in that: Also includes: Obtain a test set; wherein the test set includes test posture samples and corresponding labels; Construct a test positive sample based on two test posture samples with the same label, and construct a test negative sample based on two test posture samples with different labels; Determining a first test fusion distance between two test posture samples in the test positive sample, and a second test fusion distance between two test posture samples in the test negative sample; Determine a plurality of candidate distance thresholds, and determine false rejection rates and false acceptance rates corresponding to the plurality of candidate distance thresholds based on a first test fusion distance of the test positive sample and a second test fusion distance of the test negative sample; The candidate distance threshold whose false rejection rate and false acceptance rate are closest is determined as the distance threshold.

14. The posture authentication method according to claim 13, characterized in that: Determining a first test fusion distance between two test posture samples in the test positive sample and a second test fusion distance between two test posture samples in the test negative sample comprises: Inputting the video sequence of the test posture sample in the test positive sample and the video sequence of the test posture sample in the test negative sample into the video feature extraction model respectively to extract the corresponding test video features; Inputting the event stream data of the test posture sample in the test positive sample and the event stream data of the test posture sample in the test negative sample into the event feature extraction model respectively to extract the corresponding test event features; Calculating a test positive sample video feature distance between test video features of two test posture samples in the test positive sample, and a test positive sample event feature distance between test event features of two test posture samples in the test positive sample, and fusing the test positive sample video feature distance and the test positive sample event feature distance to obtain a first test fusion distance; Calculate the test negative sample video feature distance between the test video features of two test posture samples in the test negative sample, and the test negative sample event feature distance between the test event features of two test posture samples in the test negative sample, and fuse the test negative sample video feature distance and the test negative sample event feature distance to obtain a second test fusion distance.

15. The posture authentication method according to claim 1, characterized in that: Acquiring a second video sequence and a second event stream data of a posture to be authenticated, including: A second video sequence and second event stream data of the posture to be authenticated are acquired through a dynamic and active pixel visual sensor.

16. A posture authentication device, characterized in that: include: An acquisition module, used to acquire a first video sequence and first event stream data of a standard posture, and a second video sequence and second event stream data of a posture to be authenticated; A first extraction module, used for inputting the first video sequence and the second video sequence into a video feature extraction model respectively, so as to extract a first video feature of the first video sequence and a second video feature of the second video sequence; wherein the video feature extraction model is a model obtained by removing the classifier in the trained expanded three-dimensional network; A second extraction module is used to input the first event stream data and the second event stream data into an event feature extraction model respectively to extract a first event feature of the first event stream data and a second event feature of the second event stream data; wherein the event feature extraction model is a model obtained by removing a classifier in a spike neural network based on a self-attention mechanism that has been trained, and the event feature extraction model includes a tokenization structure connected in sequence and a plurality of superimposed converter structures; A fusion module, used to calculate a first feature distance between the first video feature and the second video feature, and a second feature distance between the first event feature and the second event feature, and fuse the first feature distance and the second feature distance to obtain a fusion distance; The authentication module is used to determine whether the fusion distance is less than a distance threshold; if so, the authentication is determined to be successful; if not, the authentication is determined to be unsuccessful.

17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the posture authentication method as claimed in any one of claims 1 to 15 when executing the computer program.

18. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a computer program, and when the computer program is executed, the steps of the posture authentication method according to any one of claims 1 to 15 are implemented.

19. A computer program product, characterized in that It comprises a computer program, which, when executed, implements the steps performed by the posture authentication method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Video-based random gesture authentication method and system

    CN113343198A