A cross-view gait recognition method with joint multi-view information bottleneck under view missing condition

Through the dual-branch network structure and adaptive selective feature fusion module, and combined with multiple loss functions, the performance reduction caused by the lack of view angle in cross-view gait recognition is solved, and higher recognition accuracy and robustness are achieved.

CN118135660BActive Publication Date: 2025-05-16SHANDONG UNIV +4
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410326444.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-05-16
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

The existing cross-view gait recognition method has reduced performance due to the lack of view angle during training, and cannot effectively handle recognition when the registered view angle is 0 degrees and the query view angle is 90 degrees.

Method used

A two-branch network structure is adopted, including a cross-view global extraction module and a cross-view consistency extraction module. Combined with an adaptive selective feature fusion module, it is trained through triple loss, cross-entropy loss and cross-view consistency loss to enhance the robustness of features and reduce the impact of view angle loss.

Benefits of technology

Effectively extract identity features that are consistent across perspectives, improve the model's performance under the condition of missing perspectives, and enhance the network's expressive ability in processing different input distribution situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118135660B_ABST
    Figure CN118135660B_ABST
Patent Text Reader

Abstract

The present invention relates to a cross-view gait recognition method with a joint multi-view information bottleneck under the condition of missing view, including: (1) constructing a network architecture of a feature extraction model and training it; including: constructing a cross-view global extraction module; constructing a cross-view consistency extraction module; constructing an adaptive selective feature fusion module; training the overall framework; (2) cross-view gait recognition; including: sending the gait profile sequence of the query data set into the trained cross-view gait recognition model with a joint multi-view information bottleneck under the condition of missing view to extract gait features, and completing the identity recognition of the query sample by comparing the feature similarity with the registered data set. The present invention enables the network to more effectively capture the key information in the input data, thereby improving the expression ability of the entire network in processing different input distribution scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a cross-view gait recognition method with a joint multi-view information bottleneck under a view missing condition, and belongs to the technical field of deep learning and pattern recognition. Background Art

[0002] Gait is the posture change and behavioral characteristics of the human body when walking, which is coordinated by the brain and various parts of the trunk. It is unique to each individual and can be used as a biological feature. Gait information collection methods are divided into wearable sensor technology and vision-based technology. Wearable sensors obtain gait characteristics by wearing accelerometers, pressure sensors and other devices, and are widely used in medical health and motion analysis. Vision-based technology uses optical cameras, event cameras and other sensors to collect gait information without the need for subject cooperation, and is gradually being used in areas such as personnel identity authentication and public security.

[0003] From the perspective of its representation of human features, gait recognition technology can be divided into model-based methods and appearance-based methods.

[0004] Model-based cross-view gait recognition methods use different methods to model human gait sequences and extract structures, motion parameters or other information as features for cross-view gait recognition. Although such methods are more in line with gait recognition technology in principle, there are some problems. For example, methods based on 3D modeling require a lot of computing resources, while methods based on skeleton extraction have high requirements for clear gait RGB images. Appearance-based methods use gait contours as input data and can be divided into methods based on energy-like graphs and methods based on gait contour sequences. The energy-like graph-based method superimposes the contour sequence and aggregates the spatiotemporal information of the gait into one image. The method based on the gait contour sequence uses the entire sequence as input, which contains rich spatiotemporal information. At present, cross-view gait recognition methods that use gait contour sequences as input and use deep learning methods to extract spatiotemporal information have achieved high recognition accuracy.

[0005] However, to date, the core focus of the gait recognition problem has been on achieving cross-view recognition, such as how to design the network to still perform well when the registration view is 0 degrees and the query view is 90 degrees. There is a serious problem of missing view in the real world, that is, the general assumption of gait recognition - obtaining a person's complete view at different angles may be insufficient. This raises a core problem: during the training process, the current view of some people is never obtained, that is, some view only appears in the test set but is missing in the training set. At this point, the above method will cause a significant reduction in performance due to its reliance on the assumption of complete view, and new means need to be proposed to solve this problem. Summary of the invention

[0006] In view of the shortcomings of the prior art, the present invention provides a cross-view gait recognition method with a joint multi-view information bottleneck under the condition of missing view; Summary of the invention:

[0008] A cross-view gait recognition method with a joint multi-view information bottleneck under the condition of missing view, including the construction of a cross-view global extraction module, a cross-view consistency extraction module, an adaptive selective feature fusion module, overall framework training and cross-view gait recognition under the condition of missing view.

[0009] In order to enhance the robustness of the learned features and reduce the impact of missing perspectives, a dual-branch network is adopted. Branch network one, the cross-perspective global extraction module, receives normal gait sequence input to extract discriminative gait features. Branch network two, the cross-perspective consistency extraction module, receives two inputs from different perspectives under the same identity to obtain identity information with cross-perspective consistency, thereby compensating for the problem of insignificant identity information caused by the missing intermediate perspective and enhancing the performance of the model. The adaptive selective feature fusion module mainly targets the problem of diversity in the distribution of missing perspectives, adaptively selects global features and consistency features, and dynamically adjusts and fuses them according to their correlation to obtain the robust representation of the final gait features. The training process mainly uses triplet loss, cross entropy loss and cross-perspective consistency loss to train the network.

[0010] Terminology explanation:

[0011] Representation encoder: an encoder that extracts pedestrian identity features;

[0012] The technical solution of the present invention is as follows:

[0013] A cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss includes the following steps:

[0014] (1) Construct the network architecture of the feature extraction model and perform training; including:

[0015] A. Construction of cross-view global extraction module

[0016] The cross-view global extraction module includes a multi-stage spatial feature extraction unit, a temporal feature extraction and horizontal pyramid pooling unit, and a separated fully connected network. The input gait profile sequence is subjected to multi-stage spatial feature extraction by the multi-stage spatial feature extraction unit, and then subjected to temporal feature extraction and horizontal pyramid pooling by the temporal feature extraction and horizontal pyramid pooling unit. Finally, the sequence is input into the separated fully connected network to map the features to a more discriminative space, thereby obtaining discriminative gait features.

[0017] B. Construction of cross-view consistency extraction module

[0018] The cross-view consistency extraction module is used to extract gait features with cross-view invariance. Specifically, it means: for two inputs from different viewpoints under the same identity, the cross-view consistency extraction module is used to extract features to obtain identity features shared across viewpoints.

[0019] C. Construction of Adaptive Selective Feature Fusion Module

[0020] The adaptive selective feature fusion module is used to fuse the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B;

[0021] The adaptive selective feature fusion module includes an adaptive feature selection component and a communication attention feature fusion component; the discriminative gait feature obtained in step A and the cross-view shared identity feature obtained in step B are selectively weighted fused through the adaptive feature selection component; the two features output after the selective weighted fusion are adaptively fused through the communication attention feature fusion component;

[0022] D. Overall framework training

[0023] The cross-view gait recognition model includes a cross-view global extraction module, a cross-view consistency extraction module, and an adaptive selective feature fusion module;

[0024] The discriminative pedestrian features output by the adaptive selective feature fusion module are used as the final representation of pedestrian gait features. The triplet loss, cross entropy loss and cross-view consistency loss are combined to train a cross-view gait recognition model with a joint multi-view information bottleneck under the condition of missing views.

[0025] (2) Cross-view gait recognition; including: sending the gait contour sequence of the query dataset into the trained cross-view gait recognition model with joint multi-view information bottleneck under view loss conditions to extract gait features, and completing the identity recognition of the query sample by comparing the feature similarity with the registered dataset.

[0026] Preferably, in step A, multi-stage spatial feature extraction comprises:

[0027] The multi-stage spatial feature extraction unit is an improved ResNet network; it includes sequentially connected convolutional layer Conv-1, convolutional layer Conv-2, pooling layer Maxpool-1, convolutional layer Conv-3, convolutional layer Conv-4, pooling layer Maxpool-2, convolutional layer Conv-5, and convolutional layer Conv-6; the spatial feature extraction is divided into three stages to maximize the extraction of spatial information;

[0028] Each stage consists of two 2D convolutional layers with 3×3 kernels, where the output channel sizes are 64, 128, and 256 respectively. The three stages are divided by two 2×2 pooling layers, so the length and width of each stage are half of the previous stage. The output of multi-stage spatial feature extraction is and Where s represents the length of the gait sequence, and h and w represent the height and width of the image.

[0029] Preferably, in step A, the temporal pooling and the horizontal pyramid pooling include:

[0030] The maximum pooling is performed on the time series dimension, i.e., the gait sequence length s, to obtain a maximum energy frame as the result of time series feature extraction; specifically, the maximum pooling is performed on the features x1, x2, and x3 along the time series dimension to obtain the time series feature extraction result, as shown in formula (I):

[0031] X i =Maxpooling s (x i ) (I)

[0032] Among them, Maxpooling(·) represents the maximum pooling operation; this step obtains the maximum energy frame of the gait sequence in the three stages, and

[0033] Horizontal pyramid pooling is used to map different body parts to different feature spaces; specifically:

[0034] For any input whose height h is a power of 2, it is divided into Block; get three new feature maps The corresponding formula is shown in formula (II):

[0035]

[0036] The maximum energy frame X obtained in the three stages i Perform horizontal uniform block division; then perform maximum pooling and average pooling and add them together; obtain the features after horizontal pyramid pooling

[0037] After concatenating the features of the first three different stages in the channel dimension, a one-dimensional convolution layer with a convolution kernel size of 1×1 is input to reduce the dimension of the features.

[0038] According to a preferred embodiment of the present invention, in step A, the input-separated fully connected network maps the features to a more discriminative space to obtain discriminative gait features; comprising:

[0039] The result obtained after dimensionality reduction is passed through a separated fully connected network to obtain the final output result of the cross-view global extraction module. The corresponding calculation formula is shown in formula (III):

[0040]

[0041] Preferably, in step B, extracting gait features having cross-view invariance includes:

[0042] The cross-view consistency extraction module includes two completely structurally consistent representation encoders and a mutual information estimator; the two representation encoders are used to generate representations of two views v1 and v2 of the same person; each representation encoder includes 4 basic residual modules;

[0043] The output of the fourth basic residual module is adjusted through the horizontal pyramid, pooling and fully connected layers to adjust the output shape, and regarded as the representation of the corresponding perspective, as shown in formula (IV):

[0044] X cvce =Linear(TP(Encoder(v1))) (IV)

[0045] In formula (IV), TP represents temporal pooling, Encoder represents the trained representation encoder, and Linear represents the fully connected layer, which is used to adjust the feature dimension;

[0046] Take one of the optimized representation encoder outputs of the view consistency representation z1 as the output of the cross-view consistency extraction module

[0047] The mutual information estimator is used to calculate the mutual information I(z1;z2) between the representations z1 and z2 of two perspectives v1 and v2 of the same person, that is, the identity features shared across perspectives. The Jensen-Shannon mutual information estimation is defined as shown in formula (V):

[0048]

[0049] Among them, sp(x)=log(1+exp(x)) is a soft addition function, JSD (z1; z2) represents the lower bound of the mutual information between z1 and z2, JSD (z1;z2) represents the lower bound of mutual information estimated using Jensen-Shannon divergence, represents the expectation of the joint distribution, represents the expectation of the marginal distribution, represents the upper bound of mutual information; and the function f φ : is a function with learnable parameters φ, as shown in formula (VI):

[0050] f φ (z1 (k) ,z2 (k) )=MLP φ (z1 (k) ||z2 (k) ) (VI)

[0051] Here, || refers to the concatenation operation; then, Monte Carlo sampling is combined to approximate the prior distribution and an approximate lower bound is estimated for the mutual information I(z1;z2), as shown in formula (VII):

[0052]

[0053] Among them, (z1 (k) ,z2 (k) ) is sampled from the joint distribution p(z1,z2), (z1 (k) ,z2 (p) ) are sampled independently from the marginal distributions p(z1) and p(z2).

[0054] Preferably, in step C, the discriminative gait features obtained in step A and the cross-viewpoint shared identity features obtained in step ) are fused, including:

[0055] The discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are selectively weighted fused through an adaptive feature selection component: including:

[0056] First, the output X of the cross-view global extraction module is cvce and the output X of the cross-view consistency extraction module cvge Perform stitching in the channel dimension;

[0057] Next, through the processing of multiple layers of fully connected layers and convolutional layers, a weight vector w is obtained;

[0058] Finally, the weight vector w is mapped to the interval (0, 1) through the Sigmoid function; the weight vector w is used to weighted merge features; the mathematical representation of this process is shown in formula (VIII):

[0059] w=σ(f conv (f fc (concat(X cvce ,X cvge )))) (VIII)

[0060] Among them, σ represents the Sigmoid function, f fc and f convThey represent the operations of the fully connected layer and the convolutional layer respectively, and concat represents concatenation in the channel dimension; as shown in formula (IX):

[0061]

[0062] For the main output stream X main and auxiliary stream X sub , using three convolutional layers as the mapping functions of query, key, and value respectively; the query Q, key K and value V are calculated by the following operations; as shown in formula (X), formula (XI), and formula (XII):

[0063] Q = Conv1d(X sub ) (X)

[0064] K = Conv1d(X main ) (XI)

[0065] V = Conv1d(X sub ) (XII)

[0066] Among them, Conv1d represents a one-dimensional convolution with a convolution kernel of 1×1;

[0067] Next, the attention score A is calculated as shown in formula (XIII):

[0068]

[0069] Among them, d k is the dimension of Q, and the weighted sum of the values ​​is obtained by using the attention score to obtain the weighted value AV; the weighted summed value is added to the main output stream X main The final feature f is obtained by weighted fusion through the learnable feature scaling parameter β final , as shown in formula (XIV):

[0070] f final =βAV+X main (XIV)

[0071] Among them, the output of the communication attention component Among them, c represents the dimension of the channel, which represents the feature fusion representation after dynamically adjusting the attention weight.

[0072] Preferably, in step D, the overall framework training includes:

[0073] When optimizing the cross-view consistency extraction module, the loss function of cross-view consistency As shown in formula (XVII):

[0074]

[0075] Among them, θ, Represents the encoder parameters, I θψ (z1; z2) represents the mutual information between the representations z1 and z2 generated by the two representation encoders; D SKL It means that D is observed from two perspectives together. KL (p θ (z1|v1)||p ψ (z2|v2)) and D KL (p ψ (z2|v2)||p θ The symmetric KL divergence is obtained by averaging the expected values ​​of (z1|v1)). The coefficient β defines the trade-off between compression rate and prediction performance.

[0076] When optimizing the overall network framework, the triplet loss function and the cross entropy loss function are introduced; the general form of the gait triplet loss is shown in formula (XV):

[0077]

[0078] in, is the triplet loss, and margin is a boundary value used to control the distance between positive samples and negative samples; Represents the anchor sample feature vector in the embedding space and the positive sample feature vector The Euclidean distance between Represents anchor samples in the embedding space and negative samples The Euclidean distance between

[0079] Cross entropy loss refers to the multi-classification cross entropy loss function, as shown in formula (XVI):

[0080]

[0081] in, is the cross entropy loss, N is the number of samples, C is the number of categories, and y ij is a binary indicator function, which is 1 if the sample belongs to the class and 0 otherwise. ij is the probability that the sample output by the network belongs to the category;

[0082] Preferably, in step E, cross-view gait recognition includes:

[0083] The registered dataset is input into the trained cross-view gait recognition model with joint multi-view information bottleneck under the condition of view loss, and the final feature representation F of the query dataset is output. g ;

[0084] The final feature representation f of each gait sample in the query dataset q With all the features F in the obtained registration dataset g The Euclidean distance calculation is performed, and the query sample is finally identified as the label of the feature with the smallest Euclidean distance in the registered dataset. The identity label of the query sample is output to complete cross-view gait recognition. The registered dataset refers to the dataset composed of gait samples with known labels entered in advance; the query dataset refers to the dataset composed of gait samples to be identified as test input.

[0085] The beneficial effects of the present invention are:

[0086] 1. The cross-view consistency extraction module proposed in the present invention can extract cross-view consistent identity features as feature supplements for the missing view when the view is missing.

[0087] 2. The adaptive selective feature fusion module proposed in the present invention can adaptively select and fuse the features extracted by the cross-view global module and the features extracted by the cross-view consistency module, so that the network can more effectively capture the key information in the input data, thereby improving the expression ability of the entire network in processing different input distribution scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 Schematic diagram of the overall network architecture of the cross-view gait recognition model in the present invention;

[0089] Figure 2 It is a schematic diagram of spatial feature extraction in the present invention;

[0090] Figure 3 Schematic diagram of optimization of the cross-view global extraction module in the present invention. DETAILED DESCRIPTION

[0091] The present invention will be further defined below in conjunction with the accompanying drawings and embodiments, but is not limited thereto.

[0092] Example 1

[0093] A cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss includes the following steps:

[0094] (1) Construct the network architecture of the feature extraction model and perform training; including:

[0095] A. Construction of cross-view global extraction module

[0096] The cross-view global extraction module includes a multi-stage spatial feature extraction unit, a temporal feature extraction and horizontal pyramid pooling unit, and a separated fully connected network. The input gait profile sequence is subjected to multi-stage spatial feature extraction by the multi-stage spatial feature extraction unit, and then subjected to temporal feature extraction and horizontal pyramid pooling by the temporal feature extraction and horizontal pyramid pooling unit. Finally, the sequence is input into the separated fully connected network to map the features to a more discriminative space, thereby obtaining discriminative gait features.

[0097] B. Construction of cross-view consistency extraction module

[0098] The cross-view consistency extraction module is used to extract gait features with cross-view invariance. Specifically, it means: for two inputs from different viewpoints under the same identity, the cross-view consistency extraction module is used to extract features to obtain identity features shared across viewpoints, so as to reduce the impact of view loss.

[0099] C. Construction of Adaptive Selective Feature Fusion Module

[0100] The adaptive selective feature fusion module is used to fuse the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B;

[0101] The adaptive selective feature fusion module includes an adaptive feature selection component and a communication attention feature fusion component; the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are selectively weighted fused through the adaptive feature selection component; the two features output after the selective weighted fusion are adaptively fused through the communication attention feature fusion component; by dynamically adjusting the attention weights, the network can more effectively capture the key information in the input data, thereby improving the expressiveness of the entire network in processing different input distribution scenarios.

[0102] D. Overall framework training

[0103] The overall framework of the present invention is shown in the attached Figure 1 The overall network architecture is a cross-view gait recognition model with a joint multi-view information bottleneck under the condition of view loss. The cross-view gait recognition model includes a cross-view global extraction module, a cross-view consistency extraction module, and an adaptive selective feature fusion module.

[0104] The discriminative pedestrian features output by the adaptive selective feature fusion module are used as the final representation of pedestrian gait features. The cross-view gait recognition model with a joint multi-view information bottleneck under the condition of missing views is trained by combining triplet loss, cross entropy loss and cross-view consistency loss to obtain a feature extraction model with good fitting ability and strong discriminability under the condition of missing views.

[0105] (2) Cross-view gait recognition; including: sending the gait contour sequence of the query dataset into the trained cross-view gait recognition model with joint multi-view information bottleneck under view loss conditions to extract gait features, and completing the identity recognition of the query sample by comparing the feature similarity with the registered dataset.

[0106] Example 2

[0107] The difference between the cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss described in Example 1 is that:

[0108] In step A, multi-stage spatial feature extraction includes:

[0109] The special feature of gait recognition is that in the training phase, the input is usually composed of 30 frames of images, while in the test phase, it is all frames. Therefore, the network design needs to be carried out under the premise of reducing the amount of calculation as much as possible. Considering the simplicity and efficiency of network design and the excellent performance of convolutional neural networks in image feature extraction, a fully convolutional network is used to extract the spatial features of pedestrians.

[0110] The gait series problem is different from person re-identification in that its input consists only of two-dimensional digital images composed of discrete pixels. Due to the lack of information such as color and background, it is undoubtedly a challenging task to obtain identity-related information that is invariant to the perspective from these discrete pixels. Shallow convolutional networks are usually able to capture fine-grained pixel information well, such as the texture, edges, and corners of the image. The output of deep convolution contains high-order information of the image, such as overall semantic information. Therefore, unlike the general network that directly uses the output of the last layer of convolution as the result of feature extraction, in order to retain the fine-grained features of the image, this section divides the spatial feature extraction into three stages, namely primary, intermediate, and advanced, to maximize the extraction of spatial information. The improved ResNet network consists of a series of convolutional layers and pooling layers. The multi-stage spatial feature extraction unit is an improved ResNet network; it includes sequentially connected convolutional layer Conv-1, convolutional layer Conv-2, pooling layer Maxpool-1, convolutional layer Conv-3, convolutional layer Conv-4, pooling layer Maxpool-2, convolutional layer Conv-5, and convolutional layer Conv-6; the spatial feature extraction is divided into the following steps: Figure 2 The three stages shown are used to maximize the extraction of spatial information; the specific structure and parameters are shown in Table 1.

[0111] Table 1

[0112]

[0113] Each stage consists of two 2D convolutional layers with 3×3 kernels, where the output channel sizes are 64, 128, and 256 respectively. The three stages are divided by two 2×2 pooling layers, so the length and width of each stage are half of the previous stage. The output of multi-stage spatial feature extraction is and Among them, s represents the length of the gait sequence, h and w represent the height and width of the image. The network structure designed in this way minimizes the amount of calculation while retaining the fine-grained features of the image.

[0114] In step A, temporal pooling and horizontal pyramid pooling include:

[0115] The maximum pooling is performed on the time series dimension, i.e., the gait sequence length s, to obtain a maximum energy frame as the result of time series feature extraction; specifically, the gait sequence output by the improved ResNet network in the previous step is expressed as and The maximum pooling is performed on the features x1, x2, and x3 along the time series dimension to obtain the time series feature extraction result, as shown in formula (I):

[0116] X i =Maxpooling s (x i ) (I)

[0117] Among them, Maxpooling(·) represents the maximum pooling operation; this step obtains the maximum energy frame of the gait sequence in the three stages, and

[0118] Two key problems emerged at this stage. First, the lengths of the feature vectors of images of different granularities output by the spatial feature extraction part are inconsistent; second, due to the particularity of gait recognition, the contribution of different body parts to identity information varies greatly. For example, compared with the head, the movement characteristics of the human legs often occupy a more important position in the recognition process. In order to comprehensively extract the local features of each body part and map the features to a more discriminative space, horizontal pyramid pooling (HPM) is used to map different body parts to different feature spaces; its process is attached. Figure 2 Specifically, it means:

[0119] For any input whose height h is a power of 2, it is divided into The heights of the features X1, X2, and X3 in the three stages are 64, 32, and 16 respectively. Due to the different feature granularities, the features of the three stages are divided into 5 blocks according to the height 16 of stage 3, and three new feature maps are obtained. The corresponding formula is shown in formula (II):

[0120]

[0121] The maximum energy frame X obtained in the three stages i Perform horizontal uniform block division; then perform maximum pooling and average pooling and add them together; obtain the features after horizontal pyramid pooling

[0122] After concatenating the features of the first three different stages in the channel dimension, a one-dimensional convolution layer with a convolution kernel size of 1×1 is input to reduce the dimension of the features.

[0123] In step A, the input-separated fully connected network maps the features to a more discriminative space to obtain discriminative gait features; including:

[0124] The result obtained after dimensionality reduction is passed through a separated fully connected network to obtain the final output result of the cross-view global extraction module. The corresponding calculation formula is shown in formula (III):

[0125]

[0126] In step B, gait features with cross-view invariance are extracted, including:

[0127] Assume that v1 and v2 are different perspectives of the same person, y is the person's identity label, and z is the representation corresponding to this label. An intuitive idea is that if the representation z contains all accessible information from v1 and v2, then the representation z must also contain the necessary label information. In addition, if the representation z only contains the common information of v1 and v2, then z will eliminate the details specific to the specific perspective and reduce the sensitivity of the perspective.

[0128] By decomposing the mutual information between the perspective v1 and its representation z1, we can get the formula:

[0129] I(v1;z1)=I(v1;z1|v2)+I(v2;z1)

[0130] Among them, v1 and v2 are different perspectives of the same person, and z1 is the representation of v1. If you want the representation z1 to be sufficient and meet the sufficiency. Then the goal at this time is to maximize the second term of the formula, that is, to let the representation z1 contain as much common information of the two perspectives as possible (which can be regarded as identity information). At the same time, it is hoped to minimize the mutual information between v1 and its representation z1 on the left side of the equation, so that the information contained in z1 is purer (only identity-related information). The intuitive meaning of the first term of the formula, the conditional mutual information I(v1; z1|v2), is the information unique to the perspective of v1 that is irrelevant to the perspective v2, that is, it can be regarded as redundant information irrelevant to the task. Therefore, by minimizing the first term of the formula I(v1; z1|v2), it can be ensured that the mutual information between v1 and z1 is minimized to reduce the information unique to v1, thereby obtaining a consistent representation of identity information under different perspectives.

[0131] Through the above analysis, it is found that we can ensure that z1 is sufficient to predict pedestrian y by ensuring that v1’s representation z1 is sufficient for v2, and by reducing I θ (z1; v1|v2), we can extract the representation of view consistency, i.e. identity information, by discarding irrelevant information. Therefore, we combine these two requirements with a relaxed Lagrangian objective to obtain the minimum sufficient representation z1 of v2.

[0132]

[0133] Similarly, the minimum sufficient representation z2 of v1 can be obtained.

[0134]

[0135] Among them, θ, represents the encoder parameters, λ1, λ2 represent the Lagrange multipliers introduced by constrained optimization. Define z1 and z2 above and reparameterize the Lagrange multipliers to obtain two loss functions and The upper bound of the mean, that is, the perspective consistency loss function:

[0136]

[0137] Among them, I θψ (z1;z2) represents the mutual information between the representations z1 and z2 generated by the two encoders. SKS It means that D is observed from two perspectives together. KL (p θ (z1|v1)||p ψ (z2|v2)) and D KL (p ψ (z2|v2)||pθ The symmetric KL divergence obtained by taking the expected average of (z1|v1). These two KL divergences are used to describe the similarities and differences between the potential representations generated by the two encoders under different views. In joint learning, by minimizing these two KL divergences, the representations of the two views can be made more consistent, prompting the model to learn shared and consistent representations across views. The coefficient β defines the trade-off between compression rate and prediction performance.

[0138] To achieve this goal, the CVCE network structure mainly consists of two completely structurally consistent representation encoders and a mutual information estimator. The two encoders are used to generate the representations of the two perspectives v1 and v2 of the same person mentioned above. Considering the excellent performance of the residual network in classification tasks and the limited amount of data in the gait recognition task, a shallow residual network is chosen as the basic architecture of the representation encoder. Here, some modifications are made to ResNet-18 to adapt to the gait task.

[0139] The cross-view consistency extraction module includes two representation encoders with completely consistent structures and a mutual information estimator; the two representation encoders are used to generate representations of two views v1 and v2 of the same person; each representation encoder includes 4 basic residual modules. A shallow residual network is used as the basic architecture of the representation encoder. Specifically, the encoder is constructed using 4 basic residual modules, and the network parameters of each basic residual module are shown in Table 2.

[0140] Table 2

[0141]

[0142] The output of the fourth basic residual module is adjusted through the horizontal pyramid, pooling and fully connected layers to adjust the output shape, and regarded as the representation of the corresponding perspective, as shown in formula (IV):

[0143] X cvce =Linear(TP(Encoder(v1))) (IV)

[0144] In formula (IV), TP represents temporal pooling, Encoder represents the trained representation encoder, and Linear represents the fully connected layer, which is used to adjust the feature dimension;

[0145] Take one of the optimized representation encoder outputs of the view consistency representation z1 as the output of the cross-view consistency extraction module

[0146] The mutual information estimator is used to calculate the mutual information I(z1;z2) between the representations z1 and z2 of the two views v1 and v2 of the same person, that is, the identity features shared across views, so as to constrain the encoder to generate pedestrian identity features that are consistent across views. In order to assign an approximate lower bound to the mutual information I(z1;z2), the Jensen-Shannon mutual information estimation method is used. Specifically, the Jensen-Shannon mutual information estimation is defined as shown in formula (V):

[0147]

[0148] Among them, sp(x)=log(1+exp(x)) is a soft addition function, JSD (z1; z2) represents the lower bound of the mutual information between z1 and z2, JSD (z1;z2) represents the lower bound of mutual information estimated using Jensen-Shannon divergence, represents the expectation of the joint distribution, represents the expectation of the marginal distribution, represents the upper bound of mutual information; and the function f φ : It refers to a function with a learnable parameter φ, which is implemented by a multi-layer perceptron and is used to distinguish whether an instance is sampled from a joint distribution. As shown in formula (VI):

[0149] f φ (z1 (k) ,z2 (k) )=MLP φ (z1 (k) ||z2 (k) ) (VI)

[0150] Here, || refers to the concatenation operation; then, Monte Carlo sampling is combined to approximate the prior distribution and an approximate lower bound is estimated for the mutual information I(z1;z2), as shown in formula (VII):

[0151]

[0152] Among them, (z1 (k) ,z2 (k) ) is sampled from the joint distribution p(z1,z2), (z1 (k) ,z2 (p) ) are independently sampled from the marginal distributions p(z1) and p(z2). The independent sampling in the second step is approximated by randomly shuffling the joint distribution p(z1,z2). During the optimization process, the Softplus function is used to ensure the stability of the optimization.

[0153] In step C, the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step ) are fused, including:

[0154] The discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are selectively weighted fused through an adaptive feature selection component: including:

[0155] First, the output X of the cross-view global extraction module is cvce and the output X of the cross-view consistency extraction module cvge Concatenate on the channel dimension to obtain the combined feature tensor combined.

[0156] Next, through the processing of multiple layers of fully connected layers and convolutional layers, a weight vector w is obtained;

[0157] Finally, the weight vector w is mapped to the interval (0, 1) through the Sigmoid function; the weight vector w is used to weighted merge features; the mathematical representation of this process is shown in formula (VIII):

[0158] w=σ(f conv (f fc (concat(X cvce ,X cvge )))) (VIII)

[0159] Among them, σ represents the Sigmoid function, f fc and f conv They represent the operations of the fully connected layer and the convolutional layer respectively, and concat represents concatenation in the channel dimension; as shown in formula (IX):

[0160]

[0161] Then, two complementary streams are obtained by weighted fusion as shown in the formula, X main For the main output stream, X sub Auxiliary flow X sub This is used for the calculation of the subsequent attention map. This adaptive feature selection process enables the network to flexibly select the mainstream and auxiliary streams of the dual-branch network. This flexibility provides the network with more powerful adaptability.

[0162] In order to further enhance the output features and capture the dependencies between the two complementary stream features, a communication attention feature fusion component is proposed to calculate the attention map using the correlation between the two features. main and auxiliary stream X sub, using three convolutional layers as mapping functions for query, key, and value respectively; the query Q, key K, and value V are calculated by the following operations; as shown in formula (X), formula (XI), and formula (XII):

[0163] Q = Conv1d(X sub ) (X)

[0164] K = Conv1d(X main ) (XI)

[0165] V = Conv1d(X sub ) (XII)

[0166] Among them, Conv1d represents a one-dimensional convolution with a convolution kernel of 1×1;

[0167] Next, the attention score A is calculated as shown in formula (XIII):

[0168]

[0169] Among them, d k is the dimension of Q, which indicates the scaled attention score introduced to prevent the inner product from being too large, thus causing the gradient to be too small. Use the attention score to weight the value and get the weighted value AV; add the weighted sum value to the main output stream X main The final feature f is obtained by weighted fusion through the learnable feature scaling parameter β final , as shown in formula (XIV):

[0170] f final =βAV+X main (XIV)

[0171] Among them, the output of the communication attention component Where c represents the dimension of the channel, which represents the feature fusion representation after dynamically adjusting the attention weights. The design of the adaptive feature fuser enables the network to more flexibly focus on and emphasize important information in the input features. By dynamically adjusting the attention weights, the network can more effectively capture the key information in the input data, thereby improving the expressiveness of the entire network in processing different input distribution scenarios.

[0172] In step D, the overall framework training includes:

[0173] The overall framework is as follows Figure 1When optimizing the cross-view consistency extraction module, given a pedestrian y, the goal is to design an objective function for its different perspectives v1 and v2. The purpose of this objective function is to discard redundant information as much as possible without sacrificing any label information. The loss function for cross-view consistency can be obtained As shown in formula (XVII):

[0174]

[0175] Among them, θ, Represents the encoder parameters, I θψ (z1; z2) represents the mutual information between the representations z1 and z2 generated by the two representation encoders; D SKL It means that D is observed from two perspectives together. KL (p θ (z1|v1)||p ψ (z2|v2)) and D KL (p ψ (z2|v2)||p θ The symmetric KL divergence obtained by averaging the expected values ​​of (z1|v1)) is used to describe the similarities and differences between the potential representations generated by the two encoders under different perspectives. In joint learning, by minimizing these two KL divergences, the representations of the two perspectives can be made more consistent, prompting the model to learn shared and consistent representations across perspectives. The coefficient β defines the trade-off between compression rate and prediction performance;

[0176] When optimizing the overall network framework, the triplet loss function and the cross entropy loss function are introduced; by selecting appropriate triplets including anchor samples, positive samples and negative samples, the embedding distance between different gait samples of the same pedestrian meets certain constraints. The general form of the gait triplet loss is shown in formula (XV):

[0177]

[0178] in, is the triplet loss, and margin is a boundary value used to control the distance between positive samples and negative samples; Represents the anchor sample feature vector in the embedding space and the positive sample feature vector The Euclidean distance between Represents anchor samples in the embedding space and negative samples The Euclidean distance between

[0179] The goal of triplet loss optimization is to make the distance between positive samples and anchor samples small enough, while the distance between negative samples and anchor samples large enough. Many advanced methods combine cross entropy loss and triplet loss to constrain the same features. Since the triplet loss cannot provide a global optimal constraint, the inter-class distance is sometimes smaller than the intra-class distance. A common approach is to combine the two losses to achieve better model performance. However, this approach ignores that the targets of the two losses are inconsistent in the embedding space. In order to suppress this inconsistency, the BNNeck layer is used to adjust the feature space. The specific inference process is shown in the attached figure. Figure 3 As shown:

[0180] Cross entropy loss refers to the multi-classification cross entropy loss function, as shown in formula (XVI):

[0181]

[0182] in, is the cross entropy loss, N is the number of samples, C is the number of categories, and y ij is a binary indicator function, which is 1 if the sample belongs to the class and 0 otherwise. ij is the probability that the sample output by the network belongs to the category; note that the feature f before the BNNeck layer is still used here i The purpose of introducing this process is to optimize the embedding space of features.

[0183] In summary, the total loss function of the overall framework training consists of three parts, namely gait triplet loss, multi-view consistency loss and cross entropy loss. It can be expressed as:

[0184] In step E, cross-view gait recognition includes:

[0185] The registered dataset is input into the trained cross-view gait recognition model with joint multi-view information bottleneck under the condition of view loss, and the final feature representation F of the query dataset is output. g ;

[0186] The final feature representation f of each gait sample in the query dataset q With all the features F in the obtained registration dataset g The Euclidean distance calculation is performed, and the query sample is finally identified as the label of the feature with the smallest Euclidean distance in the registered dataset. The identity label of the query sample is output to complete cross-view gait recognition. The registered dataset refers to the dataset composed of gait samples with known labels entered in advance; the query dataset refers to the dataset composed of gait samples to be identified as test input.

[0187] In this embodiment, the database used only contains binarized gait silhouette image sequences. Each sequence consists of one or more frames of silhouette images, and is uniformly cropped in the data preprocessing stage to have the same size of 64×44. The implementation stage follows the commonly used test protocol in the gait field. In the training stage, 30 frames are selected for each sample sequence of the above gait dataset. In the test stage, all frames are used for each sample sequence. The image size of all sample sequences is fixed to 64×44 during data preprocessing. In terms of loss function: the margin of the triple loss is fixed to 0.2, the weight of the triple loss is set to 1.0, and the weight of the cross entropy loss is set to 0.1. In terms of data sets: on the CASIA-B dataset, the batch size is set to 4×8 gait sequences in the training stage, that is, each batch size contains gait sequences of 8 perspectives in the perspective set (type×views) of 4 subjects. The input batch of the cross-perspective consistency extractor is the same as above, and each time a combination of different perspectives of the subject is taken. Optimization settings: SGD optimizer is used as the optimizer of the network in this chapter, momentum parameter is set to 0.9, weight decay is set to 0.0005. A total of 50k training times are performed on the CASIA-B dataset, the initial learning rate is set to 0.1, the learning rate is reduced to 0.01 at 20k training rounds, and the learning rate is reduced to 0.001 at 40k training rounds.

[0188] The data set was created by the Automation Research of the Chinese Academy of Sciences and contains gait data of 124 different pedestrians. Each pedestrian includes three types of gait sequences, namely Normal (NM), Bagging (BG) and Clothing (CL). The normal type includes 6 groups, marked as NM#1-6; the backpack type includes 2 groups, marked as BG#1-2; the clothing type includes 2 groups, marked as CL#1-2. Each group of types consists of 11 viewing angles (0°, 18°, 36°, 54°, 72°, 90°, 108°, 126°, 154°, 162°, 180°). Therefore, the CASIA-B gait database has a total of approximately 124×(6+2+2)×11=13640 gait sequences (very few sequences are missing). The experiment in this embodiment uses all samples of the first 74 people when training the network model on the CASIA-B data set, and the test set uses samples of the last 50 people in the data set. In the test set, NM-01, NM-02, NM-03, and NM-04 are taken as registration samples, NM-05 and NM-06 are taken as normal type query samples, BG-01 and BG-02 are taken as backpack type query samples, and CL-01 and CL-02 are taken as clothing type query samples. Set each batch size to 128, that is, select 8 subjects with different identities, and select 16 sequences for each subject. In order to verify the advancedness of the cross-view gait recognition method with a joint multi-view information bottleneck under the condition of missing view proposed in the present invention, the missing condition is designed to only include the first half of the view (0°, 18°, 36°, 54°, 72°). Comparison is made with the current five advanced gait recognition methods on the CASIA-B dataset, including GaitSet, GaitPart, GaitGL, GaitEdge, and GaitBase. The recognition accuracy of the training set containing only the first half of the view (0°, 18°, 36°, 54°, 72°) is shown in Table 3:

[0189] Table 3

[0190]

[0191]

[0192] It can be seen that when only the first half of the viewing angles (0°, 18°, 36°, 54°, 72°) are included, the average recognition rate of this method is ahead of other methods. And the best results are achieved in more than half of the registered viewing angles.

[0193] Similarly, the missing condition is designed to include only the two end viewing angles (0°, 180°). Comparison is made with the current five advanced step recognition methods on the CASIA-B dataset, including GaitSet, GaitPart, GaitGL, GaitEdge, and GaitBase. The recognition accuracy of the training set containing only the two end viewing angles (0°, 180°) is shown in Table 4:

[0194] Table 4

[0195]

[0196]

[0197] It can be seen from Table 4 that under the condition of extreme loss of perspective, this method is also at the leading level.

Claims

1. A cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss, characterized in that: The steps include: (1) Construct the network architecture of the feature extraction model and perform training; including: A. Construction of cross-view global extraction module The cross-view global extraction module includes a multi-stage spatial feature extraction unit, a temporal feature extraction and horizontal pyramid pooling unit, and a separated fully connected network. The input gait profile sequence is subjected to multi-stage spatial feature extraction by the multi-stage spatial feature extraction unit, and then subjected to temporal feature extraction and horizontal pyramid pooling by the temporal feature extraction and horizontal pyramid pooling unit. Finally, the sequence is input into the separated fully connected network to map the features to a more discriminative space, thereby obtaining discriminative gait features. In step A, multi-stage spatial feature extraction includes: The multi-stage spatial feature extraction unit is an improved ResNet network; it includes sequentially connected convolutional layer Conv-1, convolutional layer Conv-2, pooling layer Maxpool-1, convolutional layer Conv-3, convolutional layer Conv-4, pooling layer Maxpool-2, convolutional layer Conv-5, and convolutional layer Conv-6; the spatial feature extraction is divided into three stages to maximize the extraction of spatial information; B. Construction of cross-view consistency extraction module The cross-view consistency extraction module is used to extract gait features with cross-view invariance. Specifically, it means: for two inputs from different viewpoints under the same identity, the cross-view consistency extraction module is used to extract features to obtain identity features shared across viewpoints. The cross-view consistency extraction module includes two completely structurally consistent representation encoders and a mutual information estimator; the two representation encoders are used to generate representations of two views v1 and v2 of the same person; each representation encoder includes 4 basic residual modules; C. Construction of Adaptive Selective Feature Fusion Module The adaptive selective feature fusion module is used to fuse the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B; The adaptive selective feature fusion module includes an adaptive feature selection component and a communication attention feature fusion component; the discriminative gait feature obtained in step A and the cross-view shared identity feature obtained in step B are selectively weighted fused through the adaptive feature selection component; the two features output after the selective weighted fusion are adaptively fused through the communication attention feature fusion component; In step C, the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are fused, including: The discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are selectively weighted and fused through an adaptive feature selection component: including: First, the output X of the cross-view global extraction module is cvce and the output X of the cross-view consistency extraction module cvge Perform stitching in the channel dimension; Next, through the processing of multiple layers of fully connected layers and convolutional layers, a weight vector w is obtained; Finally, the weight vector w is mapped to the interval (0,1) through the Sigmoid function; D. Overall framework training The cross-view gait recognition model includes a cross-view global extraction module, a cross-view consistency extraction module, and an adaptive selective feature fusion module; The discriminative pedestrian features output by the adaptive selective feature fusion module are used as the final representation of pedestrian gait features. The triplet loss, cross entropy loss and cross-view consistency loss are combined to train a cross-view gait recognition model with a joint multi-view information bottleneck under the condition of missing views. (2) Cross-view gait recognition; including: sending the gait contour sequence of the query dataset into the trained cross-view gait recognition model with joint multi-view information bottleneck under view loss conditions to extract gait features, and completing the identity recognition of the query sample by comparing the feature similarity with the registered dataset.

2. According to the cross-view gait recognition method with joint multi-view information bottleneck under the condition of view loss in claim 1, it is characterized in that: Each stage consists of two 2D convolutional layers with 3×3 kernels, where the output channel sizes are 64, 128, and 256 respectively. The three stages are divided by two 2×2 pooling layers, so the length and width of each stage are half of the previous stage. The output of multi-stage spatial feature extraction is and Where s represents the length of the gait sequence, and h and w represent the height and width of the image.

3. According to the cross-view gait recognition method of joint multi-view information bottleneck under the condition of view loss in claim 2, it is characterized in that: In step A, temporal pooling and horizontal pyramid pooling include: The maximum pooling is performed on the time series dimension, i.e., the gait sequence length s, to obtain a maximum energy frame as the result of time series feature extraction; specifically, the maximum pooling is performed on the features x1, x2, and x3 along the time series dimension to obtain the time series feature extraction result, as shown in formula (I): X i =Maxpooling s (x i ) (I) Among them, Maxpooling(·) represents the maximum pooling operation; this step obtains the maximum energy frame of the gait sequence in the three stages, and Horizontal pyramid pooling is used to map different body parts to different feature spaces; specifically: For any input whose height h is a power of 2, it is divided into Block; get three new feature maps The corresponding formula is shown in formula (II): The maximum energy frame X obtained in the three stages " Perform horizontal uniform block division; then perform maximum pooling and average pooling and add them together; obtain the features after horizontal pyramid pooling After concatenating the features of the first three different stages in the channel dimension, a one-dimensional convolution layer with a convolution kernel size of 1×1 is input to reduce the dimension of the features.

4. According to the cross-view gait recognition method of joint multi-view information bottleneck under the condition of view loss 3, it is characterized in that: In step A, the input-separated fully connected network maps the features to a more discriminative space to obtain discriminative gait features; including: The result obtained after dimensionality reduction is passed through a separated fully connected network to obtain the final output result of the cross-view global extraction module. The corresponding calculation formula is shown in formula (III):

5. According to the cross-view gait recognition method with joint multi-view information bottleneck under the condition of view loss in claim 1, it is characterized in that: In step B, gait features with cross-view invariance are extracted, including: The output of the fourth basic residual module is adjusted through the horizontal pyramid, pooling and fully connected layers, and is regarded as a representation of the corresponding perspective, as shown in formula (IV): X cvce =Linear(TP(Encoder(v1))) (IV) In formula (IV), TP represents temporal pooling, Encoder represents the trained representation encoder, and Linear represents the fully connected layer, which is used to adjust the feature dimension; Take one of the optimized representation encoder outputs of the view consistency representation z1 as the output of the cross-view consistency extraction module The mutual information estimator is used to calculate the mutual information I(z1;z2) between the representations z1 and z2 of two perspectives v1 and v2 of the same person, that is, the identity features shared across perspectives. The Jensen-Shannon mutual information estimation is defined as shown in formula (V): Among them, sp(x)=log(1+exp(x)) is a soft addition function, JSD (z1; z2) represents the lower bound of the mutual information between z1 and z2, JSD (z1;z2) represents the lower bound of mutual information estimated using Jensen-Shannon divergence, represents the expectation of the joint distribution, represents the expectation of the marginal distribution, represents the upper bound of the mutual information; while the function is a function with learnable parameters φ, as shown in formula (VI): f φ (z1 (2) ,z2 (2) D=MLP φ (z1 (2) ||z2 (2) D (VI) Here, || refers to the concatenation operation; then, Monte Carlo sampling is combined to approximate the prior distribution and an approximate lower bound is estimated for the mutual information I(z1;z2), as shown in formula (VII): Among them, (z1 (2) ,z2 (2) D samples from the joint distribution p(z1,z2), (z1 (2) ,z2 (m) D is sampled independently from the marginal distributions p(z1) and p(z2).

6. According to the cross-view gait recognition method with joint multi-view information bottleneck under the condition of view loss in claim 1, it is characterized in that: In step C, the discriminative gait features obtained in step A and the cross-view shared identity features obtained in step B are fused, including: The weight vector w is used to weight the combined features; the mathematical representation of this process is shown in formula (VIII): w=σ(f conv (f fc (concat(X cvce ,X cvge )))) (VIII) Among them, σ represents the Sigmoid function, f fc and f conv They represent the operations of the fully connected layer and the convolutional layer respectively, and concat represents concatenation in the channel dimension; as shown in formula (IX): For the main output stream X ma"1 and auxiliary stream X s6x , using three convolutional layers as the mapping functions of query, key, and value respectively; the query Q, key K and value V are calculated by the following operations; as shown in formula (X), formula (XI), and formula (XII): Q=Conv1d(X s6x )(X) K=Conv1d(X ma"1 )(XI) V=Conv1d(X s6x )(XII) Among them, Conv1d represents a one-dimensional convolution with a convolution kernel of 1×1; Calculate the attention score A as shown in formula (XIII): Among them, d k is the dimension of Q, and the weighted sum of the values ​​is obtained by using the attention score to obtain the weighted value AV; the weighted summed value is added to the main output stream X ma"1 The final feature f is obtained by weighted fusion through the learnable feature scaling parameter β U"1a… , as shown in formula (XIV): f final =βAV+X main (XIV) Among them, the output of the communication attention feature fusion component Among them, c represents the dimension of the channel, which represents the feature fusion representation after dynamically adjusting the attention weight.

7. According to the cross-view gait recognition method with joint multi-view information bottleneck under the condition of view loss in claim 1, it is characterized in that: In step D, the overall framework training includes: When optimizing the cross-view consistency extraction module, the loss function of cross-view consistency As shown in formula (XVII): Among them, θ, Represents the encoder parameters, I θψ (z1; z2) represents the mutual information between the representations z1 and z2 generated by the two representation encoders; D SKL It means that D is observed from two perspectives together. KL (p θ (z1|v1)||p ψ (z2|v2)E and D KL (p ψ (z2|v2)||p θ The symmetric KL divergence is obtained by averaging the expected values ​​of (z1|v1)); the coefficient β defines the trade-off between compression rate and prediction performance.

8. The cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss according to claim 1 is characterized in that: In step D, the overall framework training includes: When optimizing the overall network framework, the triplet loss function and cross entropy loss function were introduced; The general form of triplet loss is shown in formula (XV): in, is the triplet loss, and margin is a boundary value used to control the distance between positive samples and negative samples; Represents the anchor sample feature vector in the embedding space and the positive sample feature vector The Euclidean distance between Represents anchor samples in the embedding space and negative samples The Euclidean distance between Cross entropy loss refers to the multi-classification cross entropy loss function, as shown in formula (XVI): in, is the cross entropy loss, N is the number of samples, C is the number of categories, and y "j is a binary indicator function, which is 1 if the sample belongs to the class and 0 otherwise. "j It is the probability that the sample output by the network belongs to the category.

9. A cross-view gait recognition method with a joint multi-view information bottleneck under the condition of view loss according to any one of claims 1-8, characterized in that: In step E, cross-view gait recognition includes: The registered dataset is input into the trained cross-view gait recognition model with joint multi-view information bottleneck under the condition of view loss, and the final feature representation F of the query dataset is output. g ; The final feature representation f of each gait sample in the query dataset q With all the features F in the obtained registration dataset g The Euclidean distance calculation is performed, and the query sample is finally identified as the label of the feature with the smallest Euclidean distance in the registered dataset. The identity label of the query sample is output to complete cross-view gait recognition. The registered dataset refers to the dataset composed of gait samples with known labels entered in advance; the query dataset refers to the dataset composed of gait samples to be identified as test input.

Citation Information

Patent Citations

  • Cross-view gait recognition method based on motion feature mixing and fine-grained multi-stage feature extraction

    CN116704611A

  • Cross-view gait recognition method based on adversarial domain adaptive learning

    CN116704612A