Multi-view gait recognition method based on twin mviT
By employing a multi-view gait recognition method based on Siamese MViT, and utilizing viewpoint feature transformation and flexible training, the problem of decreased accuracy caused by viewpoint shift is solved, thereby improving recognition accuracy and accelerating training speed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN NORMAL UNIV
- Filing Date
- 2022-09-22
- Publication Date
- 2026-04-28
AI Technical Summary
In multi-view gait recognition, the decrease in device accuracy due to viewpoint shift is a problem that existing technologies struggle to effectively address.
A multi-view gait recognition method based on Siamese MViT is adopted. The feature relationship between high-precision and low-precision angles is calculated by Siamese Mobile VisionTransformer (SMViT) model. Feature transformation and enhancement are performed by using view feature transformation tensor. Combined with the design concept of Siamese neural network and MViT model, flexible training is carried out to improve recognition accuracy.
It improved the accuracy of multi-view gait recognition, shortened the training time, and increased the recognition accuracy by 1 percentage point on the Chinese Academy of Sciences dataset, thus enhancing the robustness and fitting speed of the model.
Smart Images

Figure CN115641644B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of gait recognition technology, specifically a multi-view gait recognition method based on twin MViT. Background Technology
[0002] Gait recognition is an emerging biometric identification technology that aims to identify individuals by their walking posture. Compared with other biometric technologies, gait recognition has the advantages of being non-contact, long-distance, and difficult to spoof. In the field of intelligent video surveillance, it is more advantageous than image recognition.
[0003] In multi-view gait recognition tasks, due to viewpoint shift, the human walking contour overlaps with the torso during the shooting process, resulting in the loss of posture features. The recognition accuracy gradually decreases as the viewpoint shifts from 90° to both sides. Therefore, those skilled in the art have proposed a multi-view gait recognition method based on twin MViT. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a multi-view gait recognition method based on twin MViT, which solves the problem of decreased device accuracy caused by viewpoint shift in multi-view gait recognition.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a multi-view gait recognition method based on Siamese MViT, comprising the following processing steps:
[0006] S1: Data preprocessing step: The background subtraction method is used to extract human contour features from the CASIAB multi-view gait dataset. Then, a large amount of useless black background information is cropped, and the gait information of the subject is displayed in the center.
[0007] S2: Using the design concept of Siamese neural network, a Siamese Mobile VisionTransformer (SMViT) model was constructed for the field of multi-view gait recognition. Through the view feature relationship calculation part of the Siamese Mobile VisionTransformer (SMViT) model, the feature relationship between high-precision angle data and low-precision angle data was calculated to obtain the view feature transformation tensor.
[0008] S3: Perform feature transformation on the gait features of low-precision angles using the view feature transformation tensor calculated in step S2;
[0009] S4: The transformation feature tensor is obtained by training the view feature transformation model using Siamese Mobile Vision Transformer (SMViT);
[0010] S5: Flexibly train the Siamese Mobile Vision Transformer (SMViT) model built in step S2 by gradually shifting from high-precision angles to low-precision angles.
[0011] Preferably, the specific steps for obtaining the view feature transformation tensor in step S2 are as follows: input the high-precision set and the low-precision set into the constructed twin angle feature extraction module respectively, and obtain the view feature tensors of the high and low precision sets through angle feature correlation calculation.
[0012] Preferably, in step S2, after the view feature transformation tensor is obtained, a view transformation factor between two angle sets is constructed using the view feature transformation tensor. Then, the view transformation factor is used to enhance the low-precision set features to resemble high-precision features and weaken redundant features to obtain a new transformation feature tensor.
[0013] Preferably, the Siamese Mobile Vision Transformer (SMViT) model in step S2 is divided into two parts: the first part is the view feature relationship calculation part, and the second part is the view feature transformation verification part.
[0014] Preferably, in step S5, a similar angle feature transfer learning method is used to accelerate the fitting process. That is, the feature is diffused from a high-precision angle feature set to a low-precision angle feature set, and the model parameters of the high-precision view set are used as the initial weights of the model of the second-lowest precision view set. This accelerates the data fitting speed and, to a certain extent, enhances the accuracy.
[0015] This invention provides a multi-view gait recognition method based on twin MViT. It has the following beneficial effects:
[0016] 1. This invention utilizes a viewpoint transformation tensor to perform multi-angle gait feature transformation. Starting from the design perspective of Siamese neural networks, a simple viewpoint feature transformation method is designed. High-precision and low-precision sets are respectively input into a constructed Siamese angle feature extraction module. Through angle feature correlation calculation, viewpoint feature tensors of the high and low precision sets are obtained, and a viewpoint transformation factor between the two angle sets is constructed. Subsequently, the viewpoint transformation factor is used to enhance the low-precision set features to resemble high-precision features and weaken redundant features, resulting in a new transformation feature tensor. This feature tensor is regarded as the transformation result of the low-precision set features to high-precision features while retaining their main features, thereby improving the accuracy.
[0017] 2. This invention constructs a multi-view gait recognition model that combines the design concept of Siamese neural networks and MViT. By dividing the model into a view feature relationship calculation part and a view feature transformation verification part, the feature tensor is transformed into a new tensor through view transformation tensor calculation, and then the model is trained and discriminated through the second part to finally obtain the discrimination result, thereby improving the discrimination accuracy.
[0018] 3. This invention improves training effectiveness through a flexible transition between different perspectives and accelerates fitting by utilizing a similar angle feature transfer learning method. Specifically, it diffuses from a high-precision angle feature set to a low-precision angle feature set, using the model parameters from the high-precision view set as the initial weights for the next lower-precision view set. This accelerates data fitting and, to some extent, enhances accuracy. Ablation experiments on the Chinese Academy of Sciences dataset demonstrate that incorporating this gradual perspective training method significantly shortens training time and increases the final accuracy by approximately 1 percentage point compared to the original model without the gradual perspective method. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the overall structure of the model of the present invention;
[0020] Figure 2 This is a schematic diagram showing the block details of the present invention;
[0021] Figure 3 This is a schematic diagram of the model flow of the present invention;
[0022] Figure 4 This is a schematic diagram showing the maximum accuracy comparison of the present invention;
[0023] Figure 5 This is a schematic diagram showing the initial accuracy comparison of the present invention;
[0024] Figure 6 This is a schematic diagram comparing the training process from different angles according to the present invention. Detailed Implementation
[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Example:
[0027] like Figure 1-6 As shown, this embodiment of the invention provides a multi-view gait recognition method based on Siamese MViT, including the following processing steps:
[0028] S1: Data preprocessing step: The background subtraction method is used to extract human contour features from the CASIAB multi-view gait dataset. Then, a large amount of useless black background information is cropped, and the gait information of the subject is displayed in the center.
[0029] S2: Using the design concept of Siamese neural network, a Siamese Mobile VisionTransformer (SMViT) model was constructed for the field of multi-view gait recognition. Through the view feature relationship calculation part of the Siamese Mobile VisionTransformer (SMViT) model, the feature relationship between high-precision angle data and low-precision angle data was calculated to obtain the view feature transformation tensor.
[0030] S3: Perform feature transformation on the gait features of low-precision angles using the view feature transformation tensor calculated in step S2;
[0031] S4: The transformation feature tensor is obtained by training the view feature transformation model using Siamese Mobile Vision Transformer (SMViT);
[0032] S5: By progressively training the Siamese Mobile Vision Transformer (SMViT) model built in step S2 from high-precision angles to low-precision angles, the problem of low accuracy caused by the loss of a large number of features due to viewpoint shifts during multi-view gait recognition can be solved.
[0033] The specific steps for obtaining the view feature transformation tensor in step S2 are as follows: The high-precision set and the low-precision set are respectively input into the constructed twin angle feature extraction module, and the view feature tensors of the two precision sets are obtained through angle feature correlation calculation. In multi-view gait recognition tasks, due to viewpoint shift, the human walking contour overlaps with the torso during the shooting process, resulting in the loss of posture features. The recognition accuracy gradually decreases as the viewpoint shifts from 90° to both sides. To solve the proposed problem, the applicant, starting from the design perspective of twin neural networks, designed a simple view feature transformation method:
[0034] The specific process and theory are as follows: Since pedestrians exhibit the most and most distinct human posture features at 90°, resulting in the best recognition effect, 90° gait images are selected as the high-precision contour feature set, while 0°, 18°, 36°, 54°, 72°, 108°, 126°, 144°, 162°, and 180° are used as low-precision contour feature sets. The high-precision and low-precision sets are respectively input into a constructed twin angle feature extraction module, and the viewpoint feature tensors of the two precision sets are obtained through angle feature correlation calculation. A viewpoint transformation factor between the two angle sets is then constructed using this tensor. Subsequently, using the viewpoint transformation factor, high-precision feature enhancement and redundant feature reduction are applied to the low-precision set features, resulting in a new transformed feature tensor. This feature tensor is considered the transformation result of the low-precision set features into high-precision features while retaining their main features, thereby improving its accuracy.
[0035] In step S2, after obtaining the view feature transformation tensor, a view transformation factor between two angle sets is constructed using the view feature transformation tensor. Then, the view transformation factor is used to enhance the low-precision set features to resemble high-precision features and weaken redundant features, resulting in a new transformation feature tensor.
[0036] In step S2, the Siamese Mobile Vision Transformer (SMViT) model is divided into two main parts: The first part is the viewpoint feature relationship calculation module. This part utilizes the design approach of Siamese neural networks, combining the advantages of convolutional neural networks in spatial local feature extraction with the advantages of VisionTransformer in long-distance attention feature extraction to construct a viewpoint feature relationship calculation module. High- and low-precision gait images are input into the Siamese gait feature relationship calculation module, and by fusing pure convolutional channel features and convolutional plus ViT module channel features, the corresponding gait feature tensors are obtained. These are then processed by the Perspective Conversion module to calculate the viewpoint transformation tensor. The second part is the viewpoint feature transformation verification part. This part uses Mobile ViT as the main body, extracts features from the low-precision image set, controls the size of the feature tensor, and then calculates a new tensor through the viewpoint transformation tensor. This new tensor is then used for training and discrimination in the second part of the model, ultimately obtaining the discrimination result.
[0037] Step S5 utilizes a similar angle feature transfer learning method to accelerate fitting, which involves spreading from a high-precision angle feature set to a low-precision angle feature set. The model parameters of the high-precision view set are used as the initial weights of the model for the second-lowest precision view set, thereby accelerating the data fitting speed and enhancing accuracy to some extent.
[0038] like Figure 1 As shown, this application constructs a multi-view ViT gait recognition model that combines the design ideas of convolution and Siamese neural networks. The model is divided into two parts: the first part is the view feature relationship calculation part based on the Siamese neural network as the main design idea, and the second part is the view feature transformation verification part based on the Mobile ViT model.
[0039] First, in the first part, different angles are divided into high-precision and low-precision sets. Since the most accurate and stable angle currently used in gait recognition is 90°, the 90° dataset is used as the high-precision contour feature set, and 0°, 18°, 36°, 54°, 72°, 108°, 126°, 144°, 162°, and 180° are used as low-precision contour feature sets. Each angle set in both the high-precision and low-precision sets is then placed into the model's viewpoint feature relationship extraction module (i.e.,...). Figure 1 The module (CMBlock) first uses a Conv(N×N) module group and four Mobile modules based on MobileNetv2. A dual-channel design is employed: one channel emphasizes the spatial local feature extraction advantages of convolution, while the other highlights the global attention feature extraction advantages of ViT. Finally, the feature tensors obtained are fused using the feature fusion module AVG, and then passed through a Conv(1×1) module group and a global average pooling layer to obtain the final viewpoint feature tensor. Subsequently, the viewpoint feature tensors obtained from the high and low precision sets are passed through the PFC Block to establish viewpoint feature relationships, resulting in a gait feature transformation tensor between corresponding angles. At this point, this tensor has the same size as the feature tensor extracted from the viewpoint feature set by the CM Block.
[0040] Then, in the second part, the MVT Block module (similar to the ViT channel in the CM Block module) is used to extract the low-precision feature set into a feature tensor. This feature tensor is then transformed with the corresponding angle feature transformation tensor obtained in the first part using the IT Block. This process transforms the low-precision angle features into high-precision angle features by enhancing or weakening the feature points in the low-precision feature tensor, thereby compensating for the feature loss and weakening caused by viewpoint shift in low-precision angles. The model training uses Adam as the optimizer, with a learning rate set to 1×10⁻⁶. -3 Furthermore, the CrossEntropyLoss loss function is employed.
[0041] Finally, this paper proposes a progressive viewpoint training method. The idea behind this method is to use 90° as the initial high-precision feature set, and load the model feature weights of 90° into the training models of the similar 72° and 108°, so that the low-precision viewpoint can quickly capture the essence of the features of the similar high-precision viewpoint. Then, 72° and 108° are used as high-precision sets to continue transitioning to the two viewpoints respectively. Ablation experiments on the CASIA B multi-angle gait dataset of the Chinese Academy of Sciences have demonstrated that this method can flexibly transition viewpoints, greatly accelerating feature fitting efficiency and improving model robustness.
[0042] like Figure 2 As shown, the Conv module is divided into Conv(3 × 3) and Conv(1 × 1) depending on the kernel size. This module group consists of a conv layer and a BN layer, using SiLU as the activation function. The Mobile Block appropriately masks the Conv(1 × 1) at different positions in the processing to control the feature extraction effect and model depth. Finally, the MViT Block is a combination of convolution and ViT, using two symmetrical sets of Conv modules at the beginning and end to control the input and output sizes, and modeling local and global information in an input tensor with fewer parameters. Since the Transformer and Conv process different sizes, Rearrange is needed to control the format at the connection between the Transformer module and the Conv module.
[0043] Figure 2 The Attention module of the Transformer employs a multi-head attention mechanism. Each attention channel uses traditional scaled dot product attention. By constructing "Query," "Value," and "Key" vectors, the network assigns different attention to different features, which can be expressed by the following formula:
[0044]
[0045] X = softmax(f(Q, K))
[0046] Attention(X, V) = X × V
[0047] Multi-head attention mechanisms build upon traditional attention mechanisms, further refining self-attention and enhancing the model's ability to focus on different locations. Among these... As shown, each position in the computation process overlaps with positions in the previous layer, meaning the model doesn't just focus on the current input during learning; inputs at other times also influence the current result, thus creating temporal correlations between data. Then, concatenating the values of each attention channel with the We dot product can be represented as follows.
[0048]
[0049] MultiHead(Q,K,V)=Concat(head1,head2,...,head n W O
[0050] Furthermore, the MViT Block is a fusion model of convolution and Transformer. This module does not lose the patch order or the spatial order of pixels within each patch. By replacing local processing in convolution with deeper global processing, it models long-range nonlocal dependencies within an effective receptive domain, thus preserving the feature extraction advantages of both convolution and Transformer, and achieving better experimental results with fewer parameters.
[0051] Furthermore, the activation function used in this model is SiLU, a variant of the Sigmoid function and a special case of the Swish function, as shown in the following formula. The final pooling layer uses a global average pooling layer, as shown in the formula, where x represents an input matrix. w This represents the pooling layer's computation area.
[0052] SiLU(x) = x·Sigmoid(x)
[0053] Pooling(x w ) = Avg(x w )
[0054] PFC Block and IT Block are for calculating and inversely transforming view feature tensors. The former is used to calculate the similarity between feature tensors in two view feature sets and to derive the view transformation tensor between corresponding angles, as shown in formula (8). The latter is a view feature transformation formula, which aims to use the view transformation tensor to bring low-precision feature tensors closer to high-precision feature tensors in order to improve accuracy, as shown in formula (9). Here, x and y are the gait feature tensors of the high-precision set and the low-precision set, respectively, and N is the size of the high-precision set.
[0055]
[0056] IT(x, y) = y + PFC(x, y)
[0057] like Figure 4-5 As shown, experiments were conducted on the CASIA B multi-angle gait dataset from the Chinese Academy of Sciences: the model in this paper was tested at angles other than 90°, with direct training and progressive view training. It can be seen that in terms of final accuracy, the multi-view recognition accuracy using the SMViT model can reach over 90%, with an average accuracy of around 95% for each view. Moreover, the model in this paper trained with progressive view training has an accuracy improvement of about 1%, and its fitting speed is faster and its robustness is stronger.
[0058] like Figure 6 As shown, the accuracy changes during training from different angles on the CASIA B dataset of the Chinese Academy of Sciences show that the model trained by the progressive viewpoint has a significant improvement in feature fitting compared to the Base model, and also has a 1 percentage point improvement in maximum accuracy. The model fitting process is also more stable.
[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A multi-view gait recognition method based on Siamese MViT, characterized in that, The following processing steps are included: S1: Data preprocessing step: The background subtraction method is used to extract human contour features from the CASIAB multi-view gait dataset. Then, a large amount of useless black background information is cropped, and the gait information of the subject is displayed in the center. S2: Utilizing the design principles of Siamese neural networks, a Siamese Mobile Vision Transformer (SMViT) model was constructed for multi-view gait recognition. Through the viewpoint feature relationship calculation part of the Siamese Mobile Vision Transformer (SMViT) model, high-precision and low-precision sets were respectively input into the constructed Siamese angle feature extraction module. The viewpoint feature tensors of the high and low precision sets were obtained through angle feature correlation calculation. The Siamese Mobile Vision Transformer (SMViT) model includes a viewpoint feature relationship extraction module, a low-precision feature set extraction module, a viewpoint feature tensor calculation module, and an inverse transformation module. The similar angle feature transfer learning method is used to speed up the fitting process. That is, the feature is diffused from the high-precision angle feature set to the low-precision angle feature set. The model parameters of the high-precision view set are used as the initial weights of the model of the second-lowest precision view set. This can speed up the data fitting process and improve the accuracy to a certain extent. S3: After obtaining the view feature transformation tensor, construct the view transformation factor between the two angle sets through the view feature transformation tensor. Then, use the view transformation factor to enhance the low-precision set features to resemble high-precision features and weaken redundant features to obtain a new transformation feature tensor. S4: The new transformation feature tensor is identified using the Siamese Mobile Vision Transformer (SMViT) model to obtain multi-angle gait recognition results.
2. The multi-view gait recognition method based on Siamese MViT according to claim 1, characterized in that: In step S2, the Siamese Mobile Vision Transformer (SMViT) model is divided into two parts: the first part is the calculation of view feature relationships, and the second part is the verification of view feature transformation.
Citation Information
Patent Citations
Gait recognition method and device based on adversarial network, equipment and medium
CN112633222A