A gait recognition method and system based on adaptive feature fusion
Patent Information
- Application Number
- CN202410189194.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-20
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-02-20
AI Technical Summary
[0006]为克服上述现有技术的不足,本发明提供了一种基于自适应特征融合的步态识别方法及系统,提取更加完整且更好地表达人员身份的步态特征,解决经过水平分割和局部卷积提取的局部特征存在不同部位之间的联系被削弱的问题
[0029]本发明将自适应特征融合与深度卷积模块结合起来,通过全局特征和局部特征的自适应融合,提取更加完整且更好地表达人员身份的步态特征,解决经过水平分割和局部卷积提取的局部特征存在不同部位之间的联系被削弱的问题。
Smart Images

Figure CN118072387B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computer vision, pattern recognition and digital image processing, and particularly relates to a gait recognition method and system based on adaptive feature fusion. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Gait recognition is an individual identification method that identifies a pedestrian's walking pattern. Unlike other biometric technologies such as facial, iris, or fingerprint recognition, gait recognition allows for contactless, long-distance, and low-resolution identification. Because gait recognition does not require the active cooperation of the person being identified, it holds great promise for crime prevention, forensic identification, and public safety.
[0004] Since its inception, gait recognition has seen the development of numerous methods with promising results. However, existing methods still face challenges, with accuracy significantly influenced by factors such as clothing, carrying items, and viewing angle. Compared to traditional methods, deep learning-based gait recognition offers superior accuracy and complexity, leading to the emergence of many such methods in recent years. Some of these methods extract global or local features from gait silhouettes. For instance, GaitSet uses a 2D CNN to extract global features from gait sequences, GaitPart proposes local convolution to extract features from different body parts, and GaitGL proposes GLFE to extract both global and local information from gait silhouettes. Global features contain more spatiotemporal information about the overall gait, while local features focus more on the spatiotemporal information of different body parts. Both global and local information contribute significantly to the effectiveness of gait recognition; therefore, the thorough extraction of gait features is a crucial aspect of gait recognition.
[0005] However, in the process of local information extraction, most existing methods rely on the idea of local convolution to horizontally segment the feature map. This causes the gait features of different parts of the body to be concentrated in the horizontally segmented region, resulting in feature fragmentation. The feature map at the boundary of the segmented region is significantly weakened, which directly weakens the connection between different parts of the body. This greatly affects the feature representation of gait and ultimately reduces the accuracy of gait recognition. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, this invention provides a gait recognition method and system based on adaptive feature fusion, which extracts gait features that are more complete and better express the identity of the person, and solves the problem that the connection between different parts of the local features extracted by horizontal segmentation and local convolution is weakened.
[0007] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0008] The first aspect of this invention provides a gait recognition method based on adaptive feature fusion.
[0009] A gait recognition method based on adaptive feature fusion includes:
[0010] Obtain the gait image sequence of the person to be identified;
[0011] The preprocessed gait image sequence is used as input, and the trained gait recognition model is used to extract gait features. The gait recognition result is obtained by matching the gait features.
[0012] The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features.
[0013] Furthermore, the preprocessing involves performing binarization, alignment, resizing, and segmentation operations on the gait images in the sequence sequentially.
[0014] Furthermore, the gait recognition model includes a deep convolution module, an adaptive feature fusion module, a feature expansion module, and a feature matching module connected in sequence;
[0015] The deep convolutional model is used to extract feature map sequences.
[0016] Furthermore, the adaptive feature fusion module is used to extract global and local features, and to perform adaptive feature fusion on the global and local features, thereby compensating for the information loss of local features.
[0017] Furthermore, the adaptive feature fusion of global and local features involves autonomously finding the optimal fusion ratio between global and local features during training, and then weightedly fusing the global and local features using the optimal fusion ratio to obtain the initial gait features.
[0018] Furthermore, the feature expansion module adopts a two-stream structure, using a three-dimensional convolutional branch to retain more spatial and temporal information of the overall gait, and adaptively splices it with the backbone features in the channel dimension to obtain a more comprehensive gait representation.
[0019] Furthermore, the gait feature matching involves calculating the similarity between the gait features of the person to be identified and the gait feature database of known identities, and inferring the identity of the person to be identified based on the similarity.
[0020] The gait features of unknown identities are compared with the gait feature database of known identities, and identity information is matched from the gait feature database of known identities.
[0021] A second aspect of the present invention provides a gait recognition system based on adaptive feature fusion.
[0022] A gait recognition system based on adaptive feature fusion includes an acquisition unit and a recognition unit:
[0023] The acquisition unit is configured to acquire a sequence of gait images of the person to be identified;
[0024] The recognition unit is configured to: take the preprocessed gait image sequence as input, use the trained gait recognition model to extract gait features, and obtain the gait recognition result by matching the gait features;
[0025] The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features.
[0026] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a gait recognition method based on adaptive feature fusion as described in the first aspect of the present invention.
[0027] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a gait recognition method based on adaptive feature fusion as described in the first aspect of the present invention.
[0028] The above one or more technical solutions have the following beneficial effects:
[0029] This invention combines adaptive feature fusion with a deep convolution module. Through adaptive fusion of global and local features, it extracts gait features that are more complete and better express the identity of a person, thus solving the problem that the connection between different parts of the local features extracted by horizontal segmentation and local convolution is weakened.
[0030] This invention provides a gait recognition model that uses adaptive feature fusion as the main method to balance global and local gait features for gait recognition. It can solve the problem of the need for a large number of parameter adjustments in deep learning, reduce the number of parameter adjustments, and improve recognition accuracy.
[0031] The adaptive feature fusion block of the present invention can adaptively combine global features to compensate for the missing parts in local features, making the extracted gait features more comprehensive.
[0032] The feature extension module of this invention effectively balances the proportion of global features and local features to obtain a more comprehensive gait representation; it can solve the feature loss problem of local convolution methods and improve detection accuracy.
[0033] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0034] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0035] Figure 1 This is a flowchart of the method in the first embodiment;
[0036] Figure 2 This is a structural diagram of the gait recognition model in the first embodiment;
[0037] Figure 3 This is a diagram of the adaptive feature fusion block structure of the first embodiment;
[0038] Figure 4 This is a comparison diagram of the adaptive feature fusion block of the first embodiment and the gait features extracted by other methods.
[0039] Figure 5 This is a structural diagram of the feature extension module of the first embodiment;
[0040] Figure 6 This is a flowchart of the training process for the gait recognition model in the first embodiment; Detailed Implementation
[0041] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0042] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0043] Example 1
[0044] One embodiment of this disclosure provides a gait recognition method based on adaptive feature fusion, such as... Figure 1 As shown, it includes the following steps:
[0045] Step S1: Obtain the gait image sequence of the person to be identified.
[0046] In this embodiment, frames are continuously extracted from a single person's walking gait video to form a gait image sequence.
[0047] The gait images in the gait image sequence are preprocessed and aligned to obtain a binarized gait image sequence, specifically as follows:
[0048] (1) The contour image of a person is obtained by using the grayscale difference between corresponding pixels in the image and the background image. This contour image is not a complete outline of a person, but a rectangular image containing the outline of a person. The purpose of this is to ensure that the contour image contains the gait outline while reducing the background (black pixel area), thereby reducing the computational load of the model. The final result is a rectangular image with the gait outline as the central content and other areas as black pixels. If the grayscale value of a pixel in the current image is significantly different from that of a pixel in the background image, it is assumed that a person is walking at this pixel. A difference image is obtained by subtracting the grayscale value of the background from the grayscale value of each frame image.
[0049] (2) Perform image binarization based on the difference image:
[0050] First, each frame in the gait image sequence is converted from a color image to a grayscale image. Then, a threshold is selected to divide the pixels in the grayscale image into two categories: pixels larger than the threshold are set to white (255), and pixels smaller than the threshold are set to black (0). The final result is a black and white image where the gait outline is white and the background is black. All the images that have undergone binarization are then combined back into a sequence to obtain a binarized gait image sequence.
[0051] (3) Align the obtained binarized gait image sequence. For the contour map, find the top and bottom edges of the gait image based on the principle that the sum of the pixels in each row is not zero. Cut the contour map according to the top and bottom edges. The aligned gait image sequence contains less background information, which can accelerate the model calculation in subsequent operations and improve the model's ability to fit gait features.
[0052] (4) Resize the cut image to a height of 64 while maintaining the aspect ratio in width;
[0053] (5) Based on the principle that the maximum sum of each column is the center line, find the center, and cut the left and right sides by 22 pixels each, padding with 0 if necessary, to obtain a binary gait sequence image of size 64×44.
[0054] Step S2: Take the preprocessed gait image sequence as input, use the trained gait recognition model to extract gait features, and obtain the gait recognition result by matching the gait features;
[0055] The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features.
[0056] As can be seen, the gait recognition model that uses adaptive feature fusion as the main means and balances global and local gait features is the key part of this embodiment. The implementation process of the gait recognition model will be described in detail below.
[0057] The gait recognition model, such as Figure 2 As shown, it includes a deep convolution module, an adaptive feature fusion module, a feature expansion module, and a feature matching module connected in sequence, which will be described separately below.
[0058] I. Depthwise Convolution Module
[0059] The deep convolution module operates using a 3D convolution block. The 3D convolution block takes a gait image sequence as input, performs 3D convolution on each gait image in the sequence, extracts feature maps, and forms a feature map sequence.
[0060] 3D convolution considers image information simultaneously in terms of time, height, and width, which helps capture spatiotemporal features in gait sequences. Specifically, the operation of 3D convolution is represented as follows:
[0061] (1) Time dimension: During the convolution process, considering the time relationship between different frames in the gait image sequence helps to identify dynamic features in the gait, such as changes in steps and movement trajectories.
[0062] (2) Height and width dimensions: During the convolution process, the relationship between pixels in each gait image is considered, which helps to capture static features such as the shape and texture information of the contour.
[0063] The 3D convolutional block operation has been completed, and the resulting feature map sequence can be used for subsequent adaptive feature fusion tasks.
[0064] II. Adaptive Feature Fusion Module
[0065] like Figure 2 As shown, the adaptive feature fusion module in this embodiment uses two adaptive feature fusion blocks, with a structure of "adaptive feature fusion block - max pooling - adaptive feature fusion block".
[0066] Max pooling reduces the size of the feature map while retaining the most salient features, thereby improving the model's performance and generalization ability, and reducing the computational cost. The specific operation of max pooling is as follows:
[0067] (1) Input: The feature map sequence output by the first layer adaptive feature fusion block is used as the input of max pooling, denoted as (C,S,H,W), where C is the number of channels, S is the number of frames, H is the height, and W is the width.
[0068] (2) Pooling window: Define a fixed-size sliding window, usually a rectangular area. In this model, the pooling window size is set to 1×2×2.
[0069] (3) Sliding window: Starting from the upper left corner of the input feature map, the sliding window slides along the feature map with a certain stride to cover the entire feature map area in turn. In this model, the pooling stride is set to 1×2×2.
[0070] (4) Pooling operation: Within each window, the maximum value of the feature map within the window is selected as the output. This means that only the most salient feature within the window is retained, while other values are ignored.
[0071] (5) Output: A new feature map is generated by operating on the input feature map through a sliding window, where the maximum value of each window is retained and other values are discarded.
[0072] Among them, such as Figure 3 As shown, the adaptive feature fusion block consists of a local feature extraction block and a global feature extraction block. The specific operation is as follows:
[0073] The feature map sequences obtained by the deep convolution module are input into the local feature extraction block and the global feature extraction block, respectively, to extract local and global features.
[0074] In the local feature extraction block, the idea of horizontal segmentation is used to perform 3D convolution on different parts of the feature map and then concatenate them to obtain local information as local features; the 3D convolutions in the local feature extraction block share the same weights.
[0075] In the global feature extraction block, three-dimensional convolution is used to extract global information from the feature map sequence as global features.
[0076] Assume the extracted feature map sequence is Where C inLet S be the number of channels, S be the length of the feature map sequence, and (H, W) be the size of each frame of the feature map. To facilitate the representation of local feature extraction blocks, the input level of each frame's feature map is divided into n parts, represented as follows: The 3D convolution of the global feature extraction block is denoted as The 3D convolution of local feature extraction blocks is denoted as So, the extracted global feature Y global and local features Y local It can be represented as:
[0077]
[0078]
[0079] Among them, C out This is the number of output channels of the depthwise convolution module.
[0080] Due to the influence of horizontal segmentation, local information extracted from local feature extraction blocks is lost, and the connections between features in different parts of the gait weaken, resulting in information loss in the final extracted gait features. To compensate for this loss, global information extracted from global feature extraction blocks is used to adaptively compensate for local information. During training, the optimal fusion ratio between global and local features is autonomously found. Based on the optimal fusion ratio, global and local features are weighted and fused to output gait feature Y. AFFB Expressed as a formula:
[0081]
[0082] Among them, Y AFFB For gait features, w AFFB The optimal fusion ratio is a trainable adaptive parameter.
[0083] The output of each adaptive feature fusion block is Y. AFFB To distinguish the output of the adaptive feature fusion block from the output of the adaptive feature fusion module, the output of the entire adaptive feature fusion module, i.e., the output Y of the second adaptive feature fusion block, is used. AFFB Defined as initial gait feature V whole .
[0084] To verify the effectiveness of the adaptive feature fusion block, gait feature maps extracted by other methods were compared on the CAISA-B dataset, such as... Figure 4As shown, the first row contains gait feature maps extracted using ordinary convolution. These feature maps are comprehensive and form a cohesive whole, without any feature fragmentation. The second row contains gait feature maps extracted using local convolution. These feature maps show obvious fragmentation, weakening the connections between different body parts and losing the features at the block boundaries. The third row shows gait feature maps extracted using a combination of ordinary and local convolution. While this somewhat compensates for the loss of local features, fragmentation remains significant, and the gait features are not comprehensive. The last row contains features extracted using an adaptive feature extractor. This method does not exhibit fragmentation and, compared to feature maps extracted using ordinary convolution, enhances the features of each part.
[0085] III. Feature Extension Module
[0086] The gait feature map extracted by the adaptive feature fusion module is expanded in the channel dimension, and global and local information are rebalanced. The feature expansion module adopts a two-stream structure, using a three-dimensional convolutional branch to retain more spatial and temporal information of the overall gait, and adaptively concatenates it with the backbone features in the channel dimension to obtain a more comprehensive gait representation. The aim is to expand the initial gait features to more scales, thereby including richer gait information in the feature space. Figure 5 As shown, the dual-stream structure consists of two parts: the first part is a local feature expander, and the other part is a global feature expander. The specific process is as follows:
[0087] (1) The global feature expander uses two parallel 3D convolutions to process the input and concatenates the results in the H dimension.
[0088] (2) The local feature expander uses two parallel three-dimensional convolutional blocks to extract features from each part of the feature map after horizontal partitioning of the feature map, and then concatenates the results in the H dimension.
[0089] (3) The output of the global feature expander is multiplied by the trainable expansion weights w FEM The output of the local feature expander is concatenated with the output of the local feature expander in the channel dimension to form gait features.
[0090] The input to the feature expansion module is the output of the adaptive feature fusion module AFFB. Where C in Here, S is the number of input channels for the FEM, S is the length of the feature map sequence, and (H,W) is the size of each frame of the feature map. Each frame of the input feature map is horizontally divided into n parts, represented as follows:
[0091] Set the 3D convolution in the local feature expander to... and Set the 3D convolution in the global feature expander to... and The output Y of the global feature expander GE It can be represented as:
[0092]
[0093] The output Y of the local feature expander LE It can be represented as:
[0094]
[0095]
[0096]
[0097] Among them, C out It represents the number of output channels for a 3D convolution.
[0098] Based on the expressions for the global and local feature expanders mentioned above, the output of the feature expansion module, i.e., the gait feature Y, is... FEM , can be represented as:
[0099]
[0100] Among them, w FEM To expand the weights, they are trainable adaptive parameters.
[0101] IV. Feature Matching Module
[0102] Gait features are mapped to the feature space by adaptive horizontal pooling, and then the final gait features are obtained through a fully connected layer. Inference and matching are then performed to complete the entire process of gait recognition.
[0103] The inference and matching process involves calculating the similarity between the gait features of the person to be identified and the gait feature database of known identities, and then inferring the identity of the person to be identified based on the similarity.
[0104] The optimal fusion ratio w in gait recognition models AFFB and extended weight w FEM All of these are based on a training dataset, using cross-entropy loss and triplet loss as loss functions, and are obtained through model training. The training dataset consists of gait images of the subjects and corresponding labels, including subject number, shooting angle, etc. The training process is as follows: Figure 6 As shown.
[0105] This embodiment uses an adaptive feature fusion-based gait recognition model to solve the gait recognition problem. Through the adaptive feature fusion module, complete global features are used to supplement the lost local features, thereby significantly reducing feature loss caused by horizontal partitioning of feature maps. In addition, a feature expansion module is introduced to enrich the temporal information of gait features and adaptively balance the relationship between the detailed body information extracted by the model and the overall body information, thereby improving the gait recognition accuracy of the model under different perspectives and different walking conditions.
[0106] Example 2
[0107] One embodiment of this disclosure provides a gait recognition system based on adaptive feature fusion, including an acquisition unit and a recognition unit:
[0108] The acquisition unit is configured to acquire a sequence of gait images of the person to be identified;
[0109] The recognition unit is configured to: take the preprocessed gait image sequence as input, use the trained gait recognition model to extract gait features, and obtain the gait recognition result by matching the gait features;
[0110] The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features.
[0111] Example 3
[0112] The purpose of this embodiment is to provide a computer-readable storage medium.
[0113] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a gait recognition method based on adaptive feature fusion as described in Embodiment 1 of this disclosure.
[0114] Example 4
[0115] The purpose of this embodiment is to provide an electronic device.
[0116] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in a gait recognition method based on adaptive feature fusion as described in Embodiment 1 of this disclosure.
[0117] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A gait recognition method based on adaptive feature fusion, characterized in that, include: Obtain the gait image sequence of the person to be identified; The preprocessed gait image sequence is used as input, and the trained gait recognition model is used to extract gait features. The gait recognition result is obtained by matching the gait features. The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features. The gait recognition model includes a deep convolution module, an adaptive feature fusion module, a feature expansion module, and a feature matching module connected in sequence. The deep convolution module is used to extract feature map sequences. The adaptive feature fusion module uses two adaptive feature fusion blocks, with a structure of adaptive feature fusion block-max pooling-adaptive feature fusion block, to extract global and local features and perform adaptive feature fusion on the global and local features, thereby compensating for the information loss of local features. The adaptive feature fusion of global and local features involves finding the optimal fusion ratio between global and local features during training, and then weighting and fusing the global and local features using the optimal fusion ratio to obtain the initial gait features. The feature expansion module adopts a two-stream structure, using a three-dimensional convolutional branch to retain more spatial and temporal information of the overall gait, and adaptively splices it with the backbone features in the channel dimension to obtain a more comprehensive gait representation. The dual-stream structure consists of two parts: the first part is a local feature expander, and the second part is a global feature expander. The specific process is as follows: The global feature expander uses two parallel 3D convolutions to process the input and concatenates the results in the height dimension. The local feature expander uses two parallel 3D convolutional blocks to extract features from each part of the feature map after horizontal partitioning, and then concatenates the results in the height dimension. The output of the global feature expander is multiplied by trainable expansion weights and concatenated with the output of the local feature expander along the channel dimension to form gait features.
2. The gait recognition method based on adaptive feature fusion as described in claim 1, characterized in that, The preprocessing involves performing binarization, alignment, resizing, and cropping operations on the gait images in the sequence sequentially.
3. The gait recognition method based on adaptive feature fusion as described in claim 1, characterized in that, The gait feature matching involves calculating the similarity between the gait features of the person to be identified and the gait feature database of known identities, and inferring the identity of the person to be identified based on the similarity.
4. A gait recognition system based on adaptive feature fusion, characterized in that, The gait recognition method based on adaptive feature fusion according to any one of claims 1-3 includes an acquisition unit and a recognition unit: The acquisition unit is configured to acquire a sequence of gait images of the person to be identified; The recognition unit is configured to: take the preprocessed gait image sequence as input, use the trained gait recognition model to extract gait features, and obtain the gait recognition result by matching the gait features; The gait recognition model adaptively combines global features to compensate for the information loss of local features caused by local convolution, thereby obtaining initial gait features. The initial gait features are then expanded in the channel dimension, and global and local information are balanced again to generate the final gait features.
5. An electronic device, characterized in that it comprises: Memory is used to store computer-readable instructions in a non-transitory manner. Processor, for executing the computer-readable instructions; When the computer-readable instructions are executed by the processor, they perform the method described in any one of claims 1-3.
6. A storage medium, characterized in that, The computer-readable instructions are stored non-temporarily, wherein when the computer-readable instructions are executed by a computer, the method described in any one of claims 1-3 is performed.