A Gait Recognition Method Based on Generative Counterfactual Intervention
By introducing generative counterfactual intervention and dynamic convolution kernel generation technology into the gait recognition model, the problem of the influence of confusion factors is solved, and better interpretability and discrimination of gait representation are achieved, which significantly improves the performance of gait recognition.
Patent Information
- Application Number
- CN202310454567.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-04-25
AI Technical Summary
Existing gait recognition methods are susceptible to confusion factors during the model learning process, resulting in the inability to pay attention to the most interpretable and discriminant human areas, which in turn affects the interpretability and discriminatory nature of gait representation.
Using a generative counterfactual intervention method, the dynamic convolution kernel of matrix decomposition and diversity constraints generate facts and counterfactual features to reduce the impact of confusion factors on model learning, and improve the interpretability and discrimination of the model through the counterfactual intervention learning module.
It effectively reduces the impact of confusion factors on model learning, improves the interpretability and discriminability of gait recognition models, significantly improves the accuracy and efficiency of gait recognition, and performs well on data sets in laboratory and field scenes.
Smart Images

Figure CN116403287B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a gait recognition method based on generative counterfactual intervention. Background Art
[0002] Gait recognition is defined as the following problem: a retrieval task of identifying a specific walking pattern of a pedestrian from a video sequence of silhouette images containing the pedestrian walking and identifying the pedestrian. Gait recognition has a wide range of applications in many fields, such as smart cities, security systems, video retrieval, life and health, etc., and plays a huge role. In recent years, the main work has been to achieve better gait recognition performance by modeling the human body. However, there are often many confounding factors in the gait recognition process, resulting in the network not being able to focus on the most explanatory and discriminative human body regions. Therefore, how to eliminate the influence of confounding factors and thus model gait-related features is one of the core difficulties in gait recognition. Summary of the Invention
[0003] To solve the above problems, the purpose of the present invention is to solve the influence of confounding factors on the model learning in the gait recognition model and thus achieve an interpretable and discriminative gait representation, and provide a gait recognition method based on generative counterfactual intervention. In the present invention, first, the present invention proposes a dynamic convolution based on diversity constraints of matrix factorization as a fact / counterfactual perception kernel generator. Among them, matrix factorization ensures that the convolution kernel can efficiently achieve dynamics in a low-rank space, and the diversity constraint based on the Schatten norm can ensure the diversity of the convolution kernel from the perspective of matrix rank, thereby ensuring that the generated dynamic convolution kernel has a more powerful expression ability; in addition, this invention proposes to use generative counterfactual intervention to reduce the influence of confounding factors on model learning.
[0004] The specific technical solution adopted by the present invention is as follows:
[0005] A gait recognition method based on generative counterfactual intervention, which includes the following steps:
[0006] S1. Obtain a training data set for training a gait recognition model;
[0007] S2. Construct a gait recognition model framework, which includes a feature extraction module, a kernel generator module, a counterfactual intervention learning module, and a classifier. The input gait sequence of the gait recognition model is a set of silhouette image sequences collected during a pedestrian's walking process. The input gait sequence first passes through the feature extraction module, and a four-dimensional feature is extracted by the backbone network. Then, in the kernel generator module, a fact-aware kernel generator and a counterfactual-aware kernel generator are used to generate convolution kernels dynamically adjusted according to the input gait sequence based on the four-dimensional feature respectively. Both the fact-aware kernel generator and the counterfactual-aware kernel generator are implemented by a dynamic network based on diversity constraints and matrix factorization. Then, the convolution kernels generated by the two kernel generators are used to convolve the four-dimensional feature respectively, and the fact feature and the counterfactual feature corresponding to the input gait sequence are extracted respectively. Finally, in the counterfactual intervention learning module, the fact feature and the counterfactual feature are respectively input into the classifier, and the probability distribution of the category to which the pedestrian belongs is obtained respectively, and the triplet loss, the cross-entropy loss, and the diversity constraint are calculated. The triplet loss is calculated based on the fact feature, the cross-entropy loss is calculated by differentiating the two probability distributions output by the classifier and according to the differential result, and the diversity constraint is calculated based on the bases in the low-rank spaces of the two kernel generators.
[0008] S3. Take the weighted sum of the triplet loss, the cross-entropy loss, and the diversity constraint as the total loss function, and use the training data set to iteratively train the gait recognition model framework until convergence.
[0009] S4. When performing gait recognition, input the silhouette image sequence during the walking of the pedestrian to be queried into the trained gait recognition model framework, and perform similarity matching with the fact features corresponding to each pedestrian in the library based on the obtained fact features, and determine the pedestrian identity based on the order of the similarity values.
[0010] Preferably, the training samples included in the training data set are silhouette image sequences collected and segmented under different perspectives and different walking conditions by different people, and the scenes include laboratory scenes or outdoor scenes.
[0011] Preferably, the feature extractor adopts a 4-layer low-rank 3D CNN model, which is obtained by replacing the convolution kernels of the 3D CNN model. Among them, the full-rank 3*3*3 convolution kernels in the convolution layer of the 3D CNN model are replaced by cascaded 3*1*1 convolution kernels, 1*3*1 convolution kernels, and 1*1*3 convolution kernels.
[0012] Preferably, in the dynamic network serving as the fact-aware kernel generator and the counterfactual-aware kernel generator, the input is the four-dimensional feature X extracted from the gait sequence by the feature extraction module. The dynamic network decomposes the convolutional kernel into a static convolutional kernel W0 independent of the input and a dynamic convolutional kernel related to the input, and uses matrix factorization technology for dimensional compression. The finally output convolutional kernel is expressed as:
[0013] W(X) = W0 + PΦ(X)Q T
[0014] where P and Q are a set of bases in the low-rank space; Φ is a two-layer multi-layer perceptron for generating combination coefficients related to the input X; in the multi-layer perceptron Φ, W0, P, and Q are all learnable parameters during network training.
[0015] Preferably, the calculation formula for the diversity constraint is:
[0016]
[0017] In the formula: represents the S1 norm, P A , P C respectively represent the basis P corresponding to the fact-aware kernel generator and the basis P corresponding to the counterfactual-aware kernel generator, Q A , Q C respectively represent the basis Q corresponding to the fact-aware kernel generator and the basis Q corresponding to the counterfactual-aware kernel generator.
[0018] Preferably, the calculation formula for the cross-entropy loss is:
[0019]
[0020]
[0021] Y f = linear(X * A)
[0022] Y cf = linear(X * C)
[0023] In the formula: A represents the convolutional kernel generated by the fact-aware kernel generator, C represents the convolutional kernel generated by the counterfactual-aware kernel generator, * represents the convolution operation, linear represents the linear layer classifier, represents the differential probability distribution obtained based on the two probability distributions Y f and Y cf calculated by the linear layer classifier, Y e represents the corresponding true value.
[0024] Preferably, for generating Yf and Y cf The two linear layer classifiers share weights.
[0025] Preferably, the total loss function is the equal-weighted sum of triplet loss, cross-entropy loss, and diversity constraint.
[0026] The gait recognition method based on generative counterfactual intervention of the present invention has the following beneficial effects compared with the current gait recognition methods:
[0027] First, the gait recognition method based on generative counterfactual intervention of the present invention proposes counterfactual intervention to reduce the influence of confounding factors in model learning, thereby realizing interpretable and discriminative gait representations.
[0028] Secondly, the present invention proposes to use dynamic convolution based on matrix factorization and diversity constraint to generate factual / counterfactual features, so that these features can better describe the attributes of the samples themselves, providing sample-adaptive representations for counterfactual intervention learning.
[0029] Finally, the present invention has been greatly improved on mainstream datasets, including laboratory data (CASIA-B and OU-MVLP) and field data (Gait3D and GREW). And this method is model-independent, so it can be plugged into previous gait recognition methods.
[0030] The gait recognition method based on generative counterfactual intervention of the present invention can use counterfactual intervention and dynamic sample perception technology to reduce the influence of confounding factors on the model learning process, and can realize interpretable and discriminative gait representations. The present invention can effectively improve the accuracy and efficiency of gait recognition, and can be plugged into existing neural networks, having considerable application value. In the future, in the field of biometric recognition, the gait recognition method based on generative counterfactual intervention of the present invention can perform better in diverse complex scenarios, and can provide strong guarantees for application fields such as identity verification, park management, criminal investigation, and life and health. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 is a flowchart of the present invention;
[0032] Figure 2 is a framework diagram of the gait recognition model of the present invention;
[0033] Figure 3 is a schematic diagram of using dynamic convolution based on matrix factorization and diversity constraint of the present invention;
[0034] Figure 4 is an interpretable and discriminative attention map of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0035] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0036] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.
[0037] Reference Figure 1 , in a preferred embodiment of the present invention, a gait recognition method based on generative counterfactual intervention is provided, including the following steps:
[0038] S1. Obtain a training data set for training a gait recognition model.
[0039] In this embodiment, the training samples included in the training data set I train are silhouette image sequences collected and segmented by different persons (with unique person IDs) under different perspectives and different walking conditions, and the scenes include laboratory scenes or outdoor scenes. In the laboratory scene, it includes silhouette image sequences collected and segmented by different persons under different perspectives and different walking conditions, while in the wild scene, different pedestrians are silhouette image sequences collected and segmented under any scene perspective and any walking mode in an open scene. The above-mentioned training data set I train can be directly implemented using existing data sets such as CASIA-B (laboratory data), OU-MVLP (laboratory data), Gait3D (wild data), GREW (wild data), etc.
[0040] S2. Construct a gait recognition model framework.
[0041] Such as Figure 2As shown in the figure, the gait recognition model framework includes a feature extraction module, a kernel generator module, a counterfactual intervention learning module, and a classifier. The input gait sequence of the gait recognition model is a set of silhouette image sequences collected during a pedestrian's walking. The input gait sequence first passes through the feature extraction module, and the backbone network extracts features to form a four-dimensional feature X. Then, in the kernel generator module, a fact-aware kernel generator and a counterfactual-aware kernel generator are used to generate convolution kernels dynamically adjusted according to the input gait sequence based on the four-dimensional feature X extracted by the backbone network. Both the fact-aware kernel generator and the counterfactual-aware kernel generator are implemented by a dynamic network based on diversity constraints and matrix factorization. Then, the convolution kernels generated by the two kernel generators are used to convolve the four-dimensional feature respectively, and the fact features and counterfactual features corresponding to the input gait sequence are extracted respectively. Finally, in the counterfactual intervention learning module, the fact features and counterfactual features are input into the classifier respectively, and the probability distributions of the categories to which the pedestrians belong are obtained respectively, and the triplet loss, cross-entropy loss, and diversity constraints are calculated.
[0042] In this embodiment, the feature extractor adopts a 4-layer low-rank 3D CNN model. This low-rank 3D CNN model is constructed based on the 3D CNN model. The difference from the 3D CNN model is that the full-rank 3*3*3 convolution kernel in the convolutional layer of the 3D CNN model is replaced by three cascaded convolution kernels, which are 3*1*1 convolution kernel, 1*3*1 convolution kernel, and 1*1*3 convolution kernel in sequence.
[0043] In this embodiment, both the fact-aware kernel generator and the counterfactual-aware kernel generator adopt the aforementioned dynamic network, as Figure 3 shown. The input of this dynamic network is the four-dimensional feature X extracted by the gait sequence passing through the feature extraction module. The dynamic network decomposes the convolution kernel into a static convolution kernel W0 independent of the input and a dynamic convolution kernel related to the input, and uses matrix factorization technology for dimensional compression, and finally outputs the convolution kernel, so as to achieve the purpose of dynamic representation. The convolution kernel finally output by the dynamic network is expressed as:
[0044] W(X) = W0 + PΦ(X)Q T
[0045] where P and Q are a set of bases in the low-rank space; Φ is a two-layer multi-layer perceptron for generating combination coefficients related to the input X, where X is input into the multi-layer perceptron after average pooling. The, W0, P, and Q in the multi-layer perceptron Φ are all learnable parameters during network training and can be optimized during subsequent training.
[0046] In addition, since the dynamic convolution is generated based on low-rank factorization, the present invention proposes to use diversity constraints to enhance the representation ability of the convolution kernel, and the diversity constraints are calculated based on the bases in the low-rank spaces of the two kernel generators. Specifically, the purpose of enhancing the diversity expression is achieved by constraining the rank of the dynamic convolution kernel bases, that is, the Schatten 1-norm (S1 norm) of P and Q. The calculation formula of the final diversity constraint is:
[0047]
[0048] In the formula: represents the S1 norm, P A and P C respectively represent the basis P corresponding to the fact-aware kernel generator and the basis P corresponding to the counterfactual-aware kernel generator, Q A and Q C respectively represent the basis Q corresponding to the fact-aware kernel generator and the basis Q corresponding to the counterfactual-aware kernel generator.
[0049] In addition, after convolving the four-dimensional feature X with the convolution kernels generated by the two kernel generators respectively, the fact features and counterfactual features corresponding to the input gait sequence are extracted respectively. The triplet loss L tri can be calculated based on the fact features. The specific calculation formula of the triplet loss is a prior art and will not be described in detail here. The fact features and counterfactual features can be input into a classifier (which can be implemented by a linear layer) to obtain the probability distribution of the categories. The probability distributions corresponding to the fact features and counterfactual features can be differentiated and the cross-entropy loss can be calculated according to the differentiation result. The calculation formula of the final cross-entropy loss L cf is:
[0050]
[0051]
[0052] Y f = linear(X * A)
[0053] Y cf = linear(X * C)
[0054] In the formula: A represents the convolution kernel generated by the fact-aware kernel generator, C represents the convolution kernel generated by the counterfactual-aware kernel generator, * represents the convolution operation, linear represents the linear layer classifier, represents the differential probability distribution calculated based on the two probability distributions Y f and Y cf obtained by the linear layer classifier, Y e represents The corresponding true value. It should be noted that in this embodiment, the two linear layer classifiers used to generate Y f and Y cf share weights.
[0055] The above gait recognition model framework of the present invention can alleviate the interference of confounding factors during the learning process through counterfactual intervention learning, and express gait features by extracting interpretable and discriminative pedestrian contour edges. Among them, convolutional kernel A and convolutional kernel C are respectively generated by their corresponding perceptual generators, representing factual attention and counterfactual attention. Since counterfactuals exist with an extremely low probability, counterfactuals can cut off the correlation between attention and confounding factors, so that the correlation between model attention and model output can be analyzed. The above cross-entropy loss L cf is implemented by taking the difference between the factual / counterfactual-based outputs as the input. Since this learning process is model-independent, it can be plug-and-play in any scenario.
[0056] S3. Define the optimization objective and iteratively train the gait recognition model framework until convergence.
[0057] Take the weighted sum of the triplet loss, cross-entropy loss, and diversity constraint as the total loss function, and use the training data set to iteratively train the gait recognition model framework until convergence.
[0058] In this embodiment, the final total loss function L is the equal-weighted sum of the triplet loss, cross-entropy loss, and diversity constraint:
[0059] L = L tri + L cf + L div
[0060] The training of the model belongs to the prior art, and it can be iteratively optimized by an optimizer with the aim of minimizing the total loss function.
[0061] It should be noted that in the entire gait recognition model framework, some modules are only used during the learning process and are not required during testing and actual inference. For example, the two kernel generators and classifiers only need to use the backbone network to extract four-dimensional features, and only need to perform convolution with convolutional kernel A to output factual features for searching in the library.
[0062] S4. Perform actual gait recognition.
[0063] When performing gait recognition, input the silhouette image sequence of the pedestrian to be queried during walking into the trained gait recognition model framework, and perform similarity matching between the obtained factual features and the factual features corresponding to each pedestrian in the library, and determine the pedestrian identity based on the order of the similarity values.
[0064] Specifically, when performing actual gait recognition, only the video sequence of a pedestrian needs to be input into the trained gait recognition model, and then it can output the factual features. By comparing in the database and sorting the similarities of the features of the pedestrians most similar to the input probe to be queried, the most likely pedestrian identity can be output.
[0065] The method provided by the present invention can effectively improve the performance under various conditions. The above method will be applied to specific embodiments below so that those skilled in the art can better understand the effects of the present invention. The subsequent embodiments fully verify the effectiveness of the generative counterfactual intervention framework in the present invention with experimental results on mainstream laboratory datasets (CASIA-B and OU-MVLP) and mainstream field datasets (Gait3D and GREW).
[0066] Embodiment
[0067] The implementation method of this embodiment is as described above (i.e., S1 to S4), and the specific steps will not be elaborated in detail. Only the effects will be shown for the case data below. The present invention is implemented according to certain settings on four datasets with ground truth annotations, namely:
[0068] CASIA-B dataset: It contains 124 person IDs, 11 viewpoints, and 13,640 gait sequences. This dataset will mainly be used as a test dataset.
[0069] OU-MVLP dataset: It contains 10,307 person IDs and 14 viewpoints, and is the academic dataset with the largest target.
[0070] Gait3D and GREW datasets: They are datasets collected in an unconstrained field scenario, with characteristics such as low image quality, large pedestrian walking noise, and unfixed positions of the acquisition cameras, which are closer to the actual application scenario.
[0071] The GaitGCI method trained in this example is compared with other mainstream methods, trained and tested on the CASIA-B, OU-MVLP, Gait3D, and GREW datasets to verify the effectiveness of this method. The recognition accuracies of the final test results are shown in Tables 1, 2, 3, and 4. There is a significant improvement in the performance of the Rank-1 metric of this method on the 4 public datasets, especially in the most complex field scenario, which is more in line with the actual application scenario, laying a good foundation for real applications and model deployment, and can be well generalized to more complex location scenarios.
[0072] Table 1 Test results of this example on the CASIA-B dataset
[0073] Method NM BG CL GaitSet 95.0 87.2 70.4 GaitPart 96.2 91.5 78.7 GaitMPL 95.5 92.9 87.9 Lagrange 96.9 93.5 86.5 GaitGCI 98.4 96.6 88.5
[0074] Table 2 Test results of this example on the OU-MVLP dataset
[0075] Method Rank-1 GaitSet 87.1 GaitPart 88.7 GaitMPL 89.6 Lagrange 90.0 GaitGCI 92.1
[0076] Table 3 Test results of this example on the Gait3D dataset
[0077]
[0078] Table 4 Test results of this example on the GREW dataset
[0079] Method Rank-1 GaitSet 46.3 GaitPart 44.0 GaitGL 47.3 GaitGCI 68.5
[0080] It should be noted that in the above Table 1 and Table 2, NM represents the normal scenario, BG represents the backpack scenario, and CL represents the dressing scenario. Methods such as GaitSet, GaitPart, GLN, MT3D, CSTL, GaitGL, and 3DLocal are all advanced gait recognition models in the prior art. For details, please refer to the following prior art literature:
[0081] GaitSet:Chao,H.,He,Y.,Zhang,J.,&Feng,J.(2019,July).Gaitset:Regarding gait as a set for cross-view gait recognition.In Proceedings of the AAAI Conference on artificial intelligence(Vol.33,No.01,pp.8126-8133).
[0082] GaitPart:Fan,C.,Peng,Y.,Cao,C.,Liu,X.,Hou,S.,Chi,J.,...&He,Z.(2020).Gaitpart:Temporal part-based model for gait recognition.In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition(pp.14225-14233).
[0083] GLN: Hou, S., Cao, C., Liu, X., & Huang, Y. (2020, August). Gait lateral network: Learning discriminative and compact representations for gait recognition. In European Conference on Computer Vision (pp. 382-398). Springer, Cham.
[0084] CSTL: X. Huang et al., "Context-Sensitive Temporal Feature Learning for Gait Recognition," 2021 IEEE / CVF International Conference on Computer Vision (ICCV), 2021, pp. 12889-12898, doi:10.1109 / ICCV48922.2021.01267.
[0085] GaitGL: Lin, B., Zhang, S., & Yu, X. (2021). Gait Recognition via Effective Global-Local Feature Representation and Local Temporal Aggregation. 2021 IEEE / CVF International Conference on Computer Vision (ICCV), 14628-14636.
[0086] Lagrange Gait: Chai, T., Li, A., Zhang, S., Li, Z., & Wang, Y. (2022). Lagrange motion analysis and view embeddings for improved gait recognition. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp.
[0087] 20249-20258).
[0088] Through the above technical scheme, the present invention implements a gait recognition method based on generative counterfactuals, which uses counterfactuals to reduce the influence of confusion factors on the model learning process, thereby learning a gait representation with interpretability and discriminability, so that the model can focus on the contour area of the silhouette image representing the gait rather than the uninterpretable internal area. This model has considerable performance in mainstream laboratory scenes and field scene data sets, and is easy to implement, providing an application direction for deployment in real scenes.
[0089] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A gait recognition method based on generative counterfactual intervention, characterized in that, It includes the following steps: S1. Obtain a training data set for training a gait recognition model; S2. Construct a gait recognition model framework, the gait recognition model framework includes a feature extraction module, a kernel generator module, a counterfactual intervention learning module and a classifier. The input gait sequence of the gait recognition model is a sequence of silhouette images collected during the walking of a pedestrian. The input gait sequence first passes through the feature extraction module, and a four-dimensional feature is extracted by the backbone network; Then, in the kernel generator module, a fact-aware kernel generator and a counterfactual-aware kernel generator are used to generate convolution kernels dynamically adjusted according to the input gait sequence based on the four-dimensional feature respectively. Both the fact-aware kernel generator and the counterfactual-aware kernel generator are implemented by a dynamic network based on diversity constraints and matrix factorization. Then, the convolution kernels generated by the two kernel generators are used to perform convolution on the four-dimensional feature respectively, and the fact feature and the counterfactual feature corresponding to the input gait sequence are extracted respectively; Finally, in the counterfactual intervention learning module, the fact feature and the counterfactual feature are respectively input into the classifier to obtain the probability distribution of the category to which the pedestrian belongs, and the triplet loss, the cross-entropy loss and the diversity constraint are calculated. The triplet loss is calculated based on the fact feature. The cross-entropy loss is calculated by differentiating the two probability distributions output by the classifier and according to the differential result. The diversity constraint is calculated based on the bases in the low-rank spaces of the two kernel generators; S3. Use the weighted sum of the triplet loss, the cross-entropy loss and the diversity constraint as the total loss function, and use the training data set to iteratively train the gait recognition model framework until convergence; S4. When performing gait recognition, input the sequence of silhouette images during the walking of the pedestrian to be queried into the trained gait recognition model framework, and perform similarity matching with the fact features corresponding to each pedestrian in the library based on the obtained fact features, and determine the pedestrian identity based on the order of the similarity values.
2. The gait recognition method based on generative counterfactual intervention according to claim 1, wherein The training samples included in the training data set are sequences of silhouette images collected and segmented by different people under different perspectives and different walking conditions, and the scenes include laboratory scenes or outdoor scenes.
3. The gait recognition method based on generative counterfactual intervention according to claim 1, wherein The feature extractor uses a 4-layer low-rank 3D CNN model, and this low-rank 3D CNN model is obtained by replacing the convolution kernels of the 3D CNN model. Among them, the full-rank 3*3*3 convolution kernels in the convolution layer of the 3D CNN model are replaced by cascaded 3*1*1 convolution kernels, 1*3*1 convolution kernels and 1*1*3 convolution kernels.
4. The gait recognition method based on generative counterfactual intervention according to claim 3, wherein In the dynamic network of the fact-aware kernel generator and the counterfactual-aware kernel generator, the input is the four-dimensional feature X extracted by the gait sequence through the feature extraction module. The dynamic network decomposes the convolution kernel into a static convolution kernel W0 independent of the input and a dynamic convolution kernel related to the input, and uses matrix factorization technology for dimensional compression. The finally output convolution kernel is expressed as: W(X) = W0 + PΦ(X)Q T where P and Q are a set of bases in the low-rank space; Φ is a two-layer multi-layer perceptron for generating combination coefficients related to the input X; W0, P, and Q in the multi-layer perceptron Φ are all learnable parameters during network training.
5. The gait recognition method based on generative counterfactual intervention according to claim 4, characterized in that, The calculation formula for the diversity constraint is: In the formula: represents the S1 norm, P A , P C respectively represent the basis P corresponding to the fact-aware kernel generator and the basis P corresponding to the counterfactual-aware kernel generator, Q A , Q C respectively represent the basis Q corresponding to the fact-aware kernel generator and the basis Q corresponding to the counterfactual-aware kernel generator.
6. The gait recognition method based on generative counterfactual intervention according to claim 5, wherein The calculation formula for the cross-entropy loss is: Y f = linear(X*A) Y cf = linear(X*C) Where: A represents the convolutional kernel generated by the kernel generator of factual perception, C represents the convolutional kernel generated by the kernel generator of counterfactual perception, * represents the convolution operation, and linear represents the linear layer classifier. denotes two probability distributions Y obtained based on the linear layer classifier f and Y cf The differential probability distribution calculated, Y e represents the corresponding true value.
7. The gait recognition method based on generative counterfactual intervention according to claim 1, characterized in that, For generating Y f and Y cf The two linear layer classifiers share weights.
8. The gait recognition method based on generative counterfactual intervention according to claim 1, characterized in that, The total loss function is the weighted sum of the triplet loss, the cross-entropy loss, and the diversity constraint.
Citation Information
Patent Citations
Gait recognition method of improved loss function based on deep learning
CN111985332A
Progressive learning gait recognition method based on memory enhancement
CN114463848A