Identity recognition and authentication method for video gait recognition based on big data technology

Through a multi-scale spatiotemporal feature fusion and distributed computing framework, combined with a dynamic background suppression algorithm, the problems of low recognition rate and high resource consumption of gait recognition under low resolution and dynamic background are solved, and real-time video processing of edge devices is realized.

CN120260131AActive Publication Date: 2025-07-04郑州城发安居有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510403925.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-04
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The existing gait recognition technology has problems such as low recognition rate, high resource consumption and insufficient real-time performance in terms of low resolution, dynamic background and real-time processing of big data, making it difficult to effectively deploy on edge devices.

Method used

A multi-scale spatiotemporal feature fusion model, manifold learning dimensionality reduction, incremental learning and distributed computing framework is adopted, combined with a dynamic background suppression algorithm, and a dynamic attention mechanism is introduced through a 3D convolutional network and a bidirectional LSTM, and a joint architecture of Spark Streaming and Alluxio are used for data processing.

Benefits of technology

The recognition accuracy rate is improved by 6.4 percentage points in low resolution and dynamic background, the dynamic background segmentation error rate is reduced by 69%, the model update delay is less than 30 seconds, the throughput is improved by 75%, and the memory usage is reduced by 39.5%, which supports real-time video processing of edge computing devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260131A_ABST
    Figure CN120260131A_ABST
Patent Text Reader

Abstract

The invention discloses an identity recognition authentication method for video gait recognition based on a big data technology, and relates to the technical field of biological feature recognition, video analysis and big data processing. And carrying out identity authentication by adopting an improved cross entropy-triple mixed loss function. The method has the core advantages that an optical flow field dynamic background suppression algorithm is provided for a complex scene, under the conditions of dynamic shielding and low resolution (smaller than or equal to 640 * 480), the recognition accuracy of a CAS IA-B data set reaches 95.6% and is improved by 6.4% compared with that of a traditional ST-GCN method (89.2%), the dynamic background segmentation error rate is reduced to 4.7% from 15.2%, and the problem of feature drift is effectively solved. According to the technical scheme, the efficiency bottleneck of a traditional frame is broken through, and PB-level data processing delay lt is achieved; compared with a Hadoop scheme, the method has the advantages that the speed is increased by 75%, the memory occupation is reduced to 2.3 GB, and feasibility is provided for edge computing deployment. (character 276).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of biometric recognition, video analysis, and big data processing, and specifically provides an identity recognition and authentication method for video gait recognition based on big data technology. Background Art

[0002] Traditional gait recognition methods

[0003] Technical means:

[0004] Template matching method (such as Dynamic Time Warping DTW): relying on artificially designed gait cycle features, with a computational complexity as high as O(n) 2 ;

[0005] Optical flow method: analyzing gait through motion vectors, but with a segmentation error rate ≥ 15% under dynamic background interference (UBIRISv2 dataset);

[0006] 2D / 3D Convolutional Neural Network (such as ResNet-152): with a large number of parameters (60M), memory occupancy ≥ 3.8GB, and difficult to deploy to edge devices.

[0007] Main defects:

[0008] Poor adaptability to low resolution: when the video resolution ≤ 640×480, the recognition rate on the CASIA-B dataset drops sharply to 68%;

[0009] Loss of temporal features: traditional CNNs cannot effectively capture the long temporal dependencies of gait, resulting in a cross-view recognition rate drop ≥ 20%.

[0010] Technical bottlenecks in big data processing

[0011] Technical status quo:

[0012] Hadoop / Spark batch processing framework: it takes ≥ 3.2 hours to process 1TB of video data, unable to meet real-time requirements;

[0013] Traditional incremental learning: directly updating model parameters leads to catastrophic forgetting, with the recognition rate of old categories dropping ≥ 15% (TUM-GAID dataset).

[0014] Core problems:

[0015] Coupling of storage and computing: the storage cost of massive video data is extremely high (formula: storage volume = frame rate × resolution × duration × 3B / pixel);

[0016] Distributed synchronization delay: the time-consuming ratio of feature cross-node transmission accounts for ≥ 40%, restricting throughput (measured ≤ 1.2TB / s).

[0017] Summary of the prior art

[0018] Currently, gait recognition technology faces three major contradictions:

[0019] 1. The contradiction between accuracy and efficiency: The high accuracy of deep models comes at the cost of computational resources (e.g., ResNet-152 requires 16GB of GPU video memory);

[0020] 2. Insufficient adaptability to dynamic scenarios: Complex lighting, occlusion, and low resolution result in a recognition rate fluctuation of ≥30%;

[0021] 3. Lack of real-time performance for big data: Traditional architectures cannot achieve second-level response for PB-level data, restricting commercial implementation.

[0022] Therefore, we propose an identity recognition and authentication method for video gait recognition based on big data technology. Summary of the Invention

[0023] The purpose of the present invention is to provide an identity recognition and authentication method for video gait recognition based on big data technology.

[0024] To achieve the above purpose, the present invention provides the following technical solution: An identity recognition and authentication method for video gait recognition based on big data technology, comprising the following steps:

[0025] Step S1: Extract gait features in the video through a multi-scale spatio-temporal feature fusion model, which is composed of a cascaded 3D convolutional network and a bidirectional ConvLSTM. The kernel size of the 3D convolutional network is (3×3×3), the stride is (1×2×2), the hidden layer dimension of the bidirectional ConvLSTM is 64, and a dynamic attention mechanism is introduced, and its attention weight calculation satisfies:

[0026] ∝ t,c = Softmax(∑ x,y MLP(f t,c,x,y ));

[0027] where f is the spatio-temporal feature tensor, and t, c, x, y represent the time, channel, and spatial coordinate dimensions respectively;

[0028] Step S2: Use the manifold learning algorithm to reduce the dimension of the high-dimensional features output in Step S1, and its objective function is:

[0029]

[0030] where, is the gait cycle duration of the i-th sample,

[0031] σ = 0.5, λ = 0.01;

[0032] Step S3: Under the distributed computing framework, update the classifier parameters in an incremental learning manner. The incremental learning adopts the Elastic Weight Consolidation (EWC) algorithm, and the calculation of its importance weights satisfies:

[0033]

[0034] And it satisfies that the model update latency is less than 30 seconds, supporting real-time processing of PB-level video data;

[0035] Step S4: Perform identity classification and authentication based on an improved loss function. The loss function is a weighted combination of cross-entropy loss and triplet loss:

[0036]

[0037] As a further solution of the present invention: The step S1 further includes:

[0038] Eliminate the dynamic background interference through the background-foreground separation module, and its segmentation error rate is lower than 5%. Specifically, adopt the motion saliency detection algorithm based on the optical flow field, and its saliency threshold is set as:

[0039] T = μ + 2σ;

[0040] Where, μ and σ are respectively the mean and standard deviation of the optical flow amplitude, and the segmentation error rate on the UBIRIS v2 dataset is reduced from 15.2% to 4.7%.

[0041] As a further solution of the present invention: The distributed computing framework in the step S3 satisfies:

[0042] Adopt the joint architecture of Spark Streaming and Alluxio to achieve the decoupling of feature storage and calculation;

[0043] The data sharding granularity is adaptively adjusted according to the video resolution, and the shard size D satisfies:

[0044]

[0045] Where, R is the video resolution (unit: pixel), and F is the video frame rate.

[0046] As a further solution of the present invention: The recognition accuracy rate on the CASIA-B dataset is ≥ 95%, and the decline rate of the recognition rate in the video with a resolution lower than 640×480 is ≤ 3%.

[0047] As a further solution of the present invention: The value range of the regularization coefficient λ in the manifold learning algorithm is 0.005 ≤ λ ≤ 0.05, and the hidden layer dimension of the MLP in the dynamic attention mechanism is 128.

[0048] As a further solution of the present invention: The video gait recognition system of the method includes:

[0049] Module M1: Video preprocessing module, which performs resolution enhancement and frame rate normalization operations and outputs a video stream meeting 1080P@30fps;

[0050] Module M2: Feature extraction module, which incorporates the multi-scale spatio-temporal feature fusion model described in Claim 1;

[0051] Module M3: Distributed computing engine, which supports the Elastic Weight Consolidation (EWC) incremental learning algorithm;

[0052] Module M4: Identity authentication interface, which outputs an identification result with a confidence level ≥ 95%.

[0053] As a further solution of the present invention: The operating environment includes:

[0054] GPU cluster: No less than 8 NVIDIA V100 graphics cards, with a single-card video memory ≥ 16GB;

[0055] Distributed storage: Version above Alluxio 2.7, with a memory cache ratio ≥ 60%.

[0056] Adopting the above technical solutions, compared with the prior art, the beneficial effects of the present invention are as follows:

[0057] 1. The present invention significantly improves the recognition robustness in complex scenarios through the multi-scale spatio-temporal feature fusion model (3D CNN + dynamic attention mechanism) and the optical flow field dynamic background suppression algorithm. Under extreme conditions such as dynamic occlusion and low resolution (≤640×480), the recognition accuracy on the CASIA-B dataset reaches 95.6%, an increase of 6.4 percentage points compared with the traditional ST-GCN method (89.2%); the dynamic background segmentation error rate drops from 15.2% to 4.7%, solving the problem of feature drift caused by background interference in the prior art;

[0058] 2. The present invention realizes real-time processing of PB-level video streams through the incremental distributed learning framework (Spark + Alluxio joint architecture) and the adaptive data sharding strategy. The model update delay < 30 seconds, and the throughput reaches 2.1TB / s, which is 1.7 times faster than the Hadoop solution. The memory occupancy is only 2.3GB, a 39.5% reduction compared with similar deep learning methods (such as 3.8GB of ResNet-152), meeting the low resource requirements of edge computing devices;

[0059] 3. The present invention realizes recognition performance comparable to that of a supercomputing center on low-end hardware (an 8-node GPU cluster) through manifold learning dimensionality reduction algorithm (optimization parameter with λ = 0.01) and elastic weight consolidation (EWC) incremental learning. When the resolution is reduced to 320×240, the recognition rate only drops by 3% (while the traditional method drops by ≥25%). Meanwhile, it supports real-time parsing of 1080P@30fps videos, breaking through the bottleneck of the dependence on high-computing-power devices in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 It is a schematic diagram showing innovative focus and data support in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The following further describes the specific embodiments of the present invention with reference to the accompanying drawings. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention.

[0062] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0063] Please refer to the attached Figure 1 , an identity recognition and authentication method for video gait recognition based on big data technology according to the present invention, characterized in that it includes the following steps:

[0064] Step S1: Extract gait features in the video through a multi-scale spatio-temporal feature fusion model. The model is composed of a cascaded 3D convolutional network and a bidirectional ConvLSTM. The kernel size of the 3D convolutional network is (3×3×3), the stride is (1×2×2), the hidden layer dimension of the bidirectional ConvLSTM is 64, and a dynamic attention mechanism is introduced. The calculation of its attention weight satisfies:

[0065] ∝ t,c = Softmax(∑ x,y MLP(f t,c,x,y ));

[0066] where f is the spatio-temporal feature tensor, and t, c, x, y represent the dimensions of time, channel, and spatial coordinates respectively;

[0067] Step S2: Use the manifold learning algorithm to perform dimensionality reduction on the high-dimensional features output in Step S1. Its objective function is:

[0068]

[0069] where is the gait cycle duration of the i-th sample,

[0070] σ = 0.5, λ = 0.01;

[0071] Step S3: Under the distributed computing framework, the classifier parameters are updated in an incremental learning manner. The incremental learning adopts the Elastic Weight Consolidation (EWC) algorithm, and the calculation of its importance weights satisfies:

[0072]

[0073] And it satisfies that the model update latency is less than 30 seconds, supporting real-time processing of PB-level video data;

[0074] Step S4: Based on the improved loss function, identity classification and authentication are performed. The loss function is a weighted combination of cross-entropy loss and triplet loss:

[0075]

[0076] Example 1, Multi-scale Spatiotemporal Feature Fusion Model

[0077] Step S101: Video Input and Preprocessing

[0078] The input video resolution is standardized to 1080P@30fps. The bicubic interpolation algorithm is used to upsample the low-resolution video, and the following enhancement formula is applied:

[0079] I′ = I·G(σ = 0.5)+0.3·HistEq(I);

[0080] where G is the Gaussian filter and HistEq is the histogram equalization. After testing, the PSNR value is improved from 28.1dB to 32.3dB.

[0081] Step S102: Dynamic Attention Feature Extraction

[0082] A cascaded model of 3D CNN (kernel size 3×3×3, stride 1×2×2) and bidirectional ConvLSTM (hidden layer 64 dimensions) is constructed, and a dynamic attention mechanism is embedded:

[0083] 1 class DynamicAttention(nn.Module):

[0084] 2 def_init_(self):

[0085] 3 super0._init_0

[0086] 4 self.mlp = nn.Sequential(nn.Linear(64,128),nn.ReLU0) 5

[0088] 6 defforward(self,F):

[0089] 7#F: [Batch, Time, Channel, H, W]

[0090] 8 attn_weights = torch.softmax(self.mlp(F).mean(dim=(3, 4)), dim=1) # Weight along the time dimension

[0091] 9 return attn_weights * F

[0092] Effect verification: On the OU-MVLP dataset, introducing the attention mechanism increased the recognition rate from 85.6% to 92.3% (+6.7%), proving the innovation of spatio-temporal weight allocation

[0093] Example 2: Manifold learning dimensionality reduction and classification

[0094] Step S201: Construct the gait cycle similarity matrix

[0095] Define the similarity between samples where T i is the gait cycle duration of the i-th sample, and σ = 0.5 is selected after optimization and testing;

[0096] Comparative experiment: On the CASIA-E dataset, compared with traditional PCA dimensionality reduction:

[0097]

[0098] Step S202: Incremental classifier training

[0099] Adopt the Elastic Weight Consolidation (EWC) algorithm, and the importance weight is calculated as:

[0100]

[0101] After continuously learning 10 new classes on the TUM-GAID dataset, the recognition rate of the old classes only decreased by 1.2% (the traditional method decreased by ≥15%), proving the anti-forgetting effect.

[0102] Example 3: Distributed real-time processing system

[0103] Step S301: Adaptive data sharding

[0104] Sharding size formula:

[0105]

[0106] For example, the sharding size of a 1080P (1920×1080) @ 30fps video is:

[0107]

[0108] The measured sharding strategy enables the Spark cluster throughput to reach 2.1TB / s, which is 37% higher than the fixed sharding (512MB);

[0109] Step S302: Dynamic resource scheduling

[0110] Based on Alluxio's memory cache strategy, the cache ratio is set to ≥ 60%. When processing 1PB of video data in a 100-node cluster, traditional Hadoop takes 3.2 hours, while the present invention only takes 48 minutes, reducing latency by 75%.

[0111] Example 4: Dynamic Background Interference Suppression

[0112] Step S401: Optical flow saliency detection

[0113] Calculate the optical flow field of adjacent frames, extract the motion area, and set the significance threshold:

[0114] T = μ + 2σ;

[0115] Among them, μ and σ are the mean and standard deviation of the optical flow amplitude. On the UBIRIS v2 dataset, the dynamic background segmentation error rate dropped from 15.2% to 4.7%, and the false detection rate was reduced by 69%.

[0116] Working principle:

[0117] First, the optical flow threshold formula + bicubic enhancement is used to reduce the error rate of dynamic background segmentation in the traditional solution by 69%, and PSNR↑15%. Secondly, dynamic attention + 3DCNN / ConvLSTM cascade is used to reduce the number of parameters by 41% and improve the accuracy by 6.7% to solve the problem of serious loss of temporal features. Then, the similarity matrix S is learned using manifold ij , the recognition rate was increased by 12.6% and the time consumption was reduced by 82%. At the same time, Spark+Alluxio+EWC incremental learning increased the throughput by 75% and reduced the old class forgetting rate by 91%. Finally, the low-resolution recognition rate of the mixed loss function + confidence threshold was only less than 3%. At this point, the entire workflow is completed.

[0118] Although the present invention is disclosed as above in terms of preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modification, equivalent change and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection defined by the claims of the present invention.

Claims

1. An identity recognition and authentication method for video gait recognition based on big data technology, characterized in that: It includes the following steps: Step S1: Extract gait features in the video through a multi-scale spatio-temporal feature fusion model, which is composed of a cascaded 3D convolutional network and a bidirectional ConvLSTM. The kernel size of the 3D convolutional network is (3×3×3), the stride is (1×2×2), the hidden layer dimension of the bidirectional ConvLSTM is 64, and a dynamic attention mechanism is introduced. The calculation of its attention weight satisfies: ∝ t,c = Softmax(∑ x,y MLP(f t,c,x,y )); where f is the spatio-temporal feature tensor, and t, c, x, y represent the time, channel, and spatial coordinate dimensions respectively; Step S2: Use the manifold learning algorithm to reduce the dimension of the high-dimensional features output in Step S1, and its objective function is: Among them, is the gait cycle duration of the i-th sample, σ = 0.5, λ = 0.01; Step S3: Under the distributed computing framework, update the classifier parameters in an incremental learning manner. The incremental learning adopts the Elastic Weight Consolidation (EWC) algorithm, and the calculation of its importance weight satisfies: And it satisfies that the model update delay is less than 30 seconds, supporting real-time processing of PB-level video data; Step S4: Perform identity classification and authentication based on an improved loss function, and the loss function is a weighted combination of cross-entropy loss and triplet loss:

2. The identity recognition and authentication method for video gait recognition based on big data technology according to claim 1, characterized in that: The following is also included in Step S1: Eliminate dynamic background interference through a background-foreground separation module, and its segmentation error rate is lower than 5%. Specifically, use a motion saliency detection algorithm based on the optical flow field, and its saliency threshold is set as: T = μ + 2σ; where μ and σ are the mean and standard deviation of the optical flow amplitude respectively, and the segmentation error rate on the UBIRIS v2 dataset is reduced from 15.2% to 4.7%.

3. The identity recognition and authentication method for video gait recognition based on big data technology according to claim 1, wherein: The distributed computing framework in Step S3 satisfies: Adopt the joint architecture of Spark Streaming and Alluxio to achieve the decoupling of feature storage and calculation; The data sharding granularity is adaptively adjusted according to the video resolution, and the shard size D satisfies: where R is the video resolution (unit: pixel), and F is the video frame rate.

4. An identity recognition and authentication method for video gait recognition based on big data technology according to claim 1, characterized in that: The recognition accuracy on the CASIA-B dataset is ≥95%, and the decline rate of the recognition rate in videos with a resolution lower than 640×480 is ≤3%.

5. The identity recognition and authentication method for video gait recognition based on big data technology according to claim 1, characterized in that: The value range of the regularization coefficient λ in the manifold learning algorithm is 0.005 ≤ λ ≤ 0.05, and the hidden layer dimension of the MLP in the dynamic attention mechanism is 128.

6. A method for identity recognition and authentication of video gait recognition based on big data technology according to any one of claims 1-5, characterized in that: The video gait recognition system of the method includes: Module M1: Video preprocessing module, which performs resolution enhancement and frame rate standardization operations and outputs a video stream meeting 1080P@30fps; Module M2: Feature extraction module, which internally contains the multi-scale spatio-temporal feature fusion model described in Claim 1; Module M3: Distributed computing engine, which supports the Elastic Weight Consolidation (EWC) incremental learning algorithm; Module M4: Identity authentication interface, which outputs a recognition result with a confidence level ≥95%.

7. An identity recognition and authentication method for video gait recognition based on big data technology according to claim 6, characterized in that: The operating environment includes: GPU cluster: At least 8 NVIDIA V100 graphics cards, and the single-card video memory ≥16GB; Distributed storage: Version 2.7 or above of Alluxio, and the memory cache ratio ≥60%.

Citation Information

Patent Citations

  • Identity recognition method based on multi-channel space-time network and joint optimization loss

    CN112131970A

  • Cross-view gait recognition method based on spatio-temporal information enhancement and multi-scale saliency feature extraction

    CN113947814A

  • Depression risk assessment method and device based on multi-modal gait feature fusion

    CN117393159A

  • AI video low-altitude target identification and real-time tracking method based on deep learning

    CN119723421A

  • Automatically classifying animal behavior

    US20190087965A1