A gait recognition method, device, equipment and storage medium

The RGB image sequence is denoised through the potential diffusion model and geometric driving feature matching module to generate static and dynamic gait feature fields, which solves the problem of noise interference in the existing gait recognition methods and improves the recognition accuracy.

CN120198970BActive Publication Date: 2025-07-22SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510669567.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-07-22
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing gait recognition method cannot effectively remove identity-independent noise in video, resulting in low recognition accuracy and it is difficult for the prior art to establish a direct mapping relationship between image pixel-level information and identity features.

Method used

The pre-trained latent diffusion model is used to denoise the RGB image sequence, and static and dynamic gait feature fields are generated through intra- and inter-frame feature matching. The latent diffusion model and geometric-driven feature matching module are used to denoise, and the identity-independent redundant information is gradually stripped away.

Benefits of technology

It improves the extraction quality and recognition accuracy of gait representation, effectively removes background interference and clothing texture changes, and enhances the expression ability of gait characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198970B_ABST
    Figure CN120198970B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the field of biometric recognition technology, and provides a gait recognition method, device, equipment and storage medium. The method includes: obtaining an RGB image sequence of a pedestrian, performing denoising processing on the RGB image sequence by using a pre-trained latent diffusion model to obtain the diffusion features of each frame of RGB image in the RGB image sequence, respectively performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each diffusion feature to obtain a static gait feature field and a dynamic gait feature field of the pedestrian, wherein both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields, and recognizing the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian, so as to realize gait feature extraction through a denoising process of continuously removing identity-irrelevant features, and improve the extraction quality of gait representation and the accuracy of gait recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biometric recognition, and particularly relates to a gait recognition method, device, equipment and storage medium. Background Art

[0002] The differences among pedestrians in videos are very subtle because videos are in a very high-dimensional space and there is a lot of noise unrelated to identity, which makes it a very challenging task to extract corresponding biometric features (gait recognition) from pedestrian videos and has attracted increasing attention.

[0003] Current gait recognition methods can be mainly divided into three technical routes: unimodal gait recognition, multimodal gait recognition, and RGB-based gait recognition. Among them, unimodal gait recognition methods usually first use pre-designed networks (such as human segmentation networks, human semantic analysis networks, human key point prediction networks, etc.) to extract gait modalities related to identity, and remove noise information unrelated to identity such as background and clothing color through this explicit preprocessing. Then, the obtained gait modalities (such as silhouette maps, pose heat maps, etc.) are input into the gait recognition network for corresponding recognition. Although this separate design can effectively reduce background interference, it may cause loss of identity-related information; Multimodal gait recognition methods improve the recognition accuracy by integrating multiple complementary gait modalities (such as combining contour maps, skeleton sequences, and semantic segmentation maps). Although multi-source information fusion can theoretically enhance the feature expression ability, these methods rely on third-party pre-trained models to obtain modal features. Since these models are limited by original supervision tasks unrelated to identity (such as pose estimation, semantic segmentation), the extracted modal features still have the inherent defect of insufficient correlation with identity, and lack a unified feature space alignment mechanism, resulting in a bottleneck in improving the accuracy of multimodal fusion; For RGB-based gait recognition, there are now many RGB-based end-to-end gait recognition methods. These methods usually draw on existing gait modal features (such as silhouettes, SMPL models, human skeleton points, etc.), directly extract features from RGB images, and input them into the downstream gait recognition network for recognition. Although these methods can retain richer visual information, they require complex designs to simulate the functions of professional preprocessing networks and fail to effectively establish a direct mapping relationship between image pixel-level information and identity features, resulting in difficulty in effectively separating identity features from interference factors. Therefore, there is an urgent need for a new gait recognition method to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a gait recognition method, device, equipment and storage medium, aiming to solve the problem of low gait recognition accuracy caused by the existing technology.

[0005] In a first aspect, the present invention provides a gait recognition method, and the method includes the following steps:

[0006] Obtain an RGB image sequence of a pedestrian;

[0007] Use a pre-trained latent diffusion model to denoise the RGB image sequence, and obtain the diffusion features of each frame of RGB image in the RGB image sequence;

[0008] Perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields;

[0009] Recognize the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

[0010] In some embodiments, the step of using a pre-trained latent diffusion model to denoise the RGB image sequence includes:

[0011] Use the image encoder in the latent diffusion model to project each frame of RGB image in the RGB image sequence into the latent space to obtain corresponding latent features;

[0012] According to the denoising time step determined by training the latent diffusion model, use the U-net network in the latent diffusion model to perform one-step denoising on the latent features to obtain the diffusion features.

[0013] In some embodiments, the step of performing intra-frame feature matching and inter-frame feature matching with frame geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian includes:

[0014] According to the diffusion features and the human silhouette pre-extracted from the RGB image corresponding to the diffusion features, use a preset feature generation strategy to generate query features in the spatial dimension and key features in the temporal dimension of the RGB image corresponding to the diffusion features;

[0015] According to the key features, assign directions to each pixel in the query features to obtain the static gait feature field and the dynamic gait feature field.

[0016] In some embodiments, the step of using a preset feature generation strategy to generate query features in the spatial dimension and key features in the temporal dimension of the RGB image corresponding to the diffusion features includes:

[0017] According to the diffusion feature and the human silhouette, use the query feature calculation formula to calculate the query feature, where represents the query feature corresponding to the th frame of RGB image, represents the human silhouette of the th frame of RGB image, represents the diffusion feature of the th frame of RGB image, represents the first sub-module;

[0018] According to the preset time reception range, use the key feature calculation formula to calculate the key feature, where represents the key feature corresponding to the th frame of RGB image, represents the time reception range, represents the second sub-module.

[0019] In some embodiments, the step of performing direction assignment on each pixel in the query feature according to the key feature includes:

[0020] Obtain the set of neighboring pixels of each pixel in the query feature within the key feature;

[0021] Calculate the similarity distribution between each pixel in the query feature and its corresponding set of neighboring pixels;

[0022] Perform direction assignment on each pixel in the query feature according to the similarity distribution and a preset direction template.

[0023] In a second aspect, the present invention provides a gait recognition device, the device includes:

[0024] An image acquisition unit, configured to acquire a sequence of RGB images of a pedestrian;

[0025] An image denoising unit, configured to perform denoising processing on the sequence of RGB images by using a pre-trained latent diffusion model to obtain the diffusion feature of each frame of RGB image in the sequence of RGB images;

[0026] A feature matching unit, configured to perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features respectively to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields;

[0027] A gait recognition unit, configured to recognize the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait feature of the pedestrian.

[0028] In some embodiments, the image denoising unit includes:

[0029] An image projection unit, configured to project each frame of RGB image in the RGB image sequence into a latent space by using an image encoder in the latent diffusion model, so as to obtain corresponding latent features;

[0030] A feature denoising unit, configured to perform one-step denoising on the latent features by using a U-net network in the latent diffusion model according to a denoising time step determined by training the latent diffusion model, so as to obtain the diffusion features.

[0031] In some embodiments, the feature matching unit includes:

[0032] A feature generation unit, configured to generate a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion features by using a preset feature generation strategy according to the diffusion features and a human silhouette pre-extracted from the RGB image corresponding to the diffusion features;

[0033] A direction assignment unit, configured to perform direction assignment on each pixel in the query feature according to the key feature, so as to obtain the static gait feature field and the dynamic gait feature field.

[0034] In a third aspect, the present invention further provides a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the method described above are implemented.

[0035] In a fourth aspect, the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0036] When performing gait recognition in an embodiment of the present invention, an RGB image sequence of a pedestrian is acquired, and the RGB image sequence is denoised by using a pre-trained latent diffusion model to obtain diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian, so that gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, and the extraction quality of gait representation and the accuracy of gait recognition are improved. Description of the Drawings

[0037] Figure 1 It is a schematic flowchart of a gait recognition method provided in the first embodiment of the present invention;

[0038] Figure 2 It is a schematic structural diagram of a feature matching module in a gait recognition method provided in the first embodiment of the present invention;

[0039] Figure 3 It is a schematic structural diagram of the first / second sub-module in the feature matching module provided in the first embodiment of the present invention;

[0040] Figure 4 It is a schematic structural diagram of an improved GaitBase provided in the first embodiment of the present invention;

[0041] Figure 5 It is a schematic structural diagram of a gait recognition device provided in the second embodiment of the present invention;

[0042] Figure 6 It is a schematic structural diagram of a computing device provided in the third embodiment of the present invention. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations. And the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. The terms "first", "second" and similar terms do not denote any order, quantity or importance, but are only used to distinguish different components. "Connection" or "connected" and similar terms are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. The term "plural" means two or more, and other quantifiers are similar.

[0045] To keep the following description of the embodiments of the present invention clear and concise, the detailed descriptions of some known functions and known components are omitted in this specification.

[0046] The following describes the specific implementation of the present invention in detail with reference to specific embodiments:

[0047] Embodiment 1:

[0048] Figure 1 The implementation process of the gait recognition method provided in Embodiment 1 of the present invention is shown. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown and are described in detail as follows:

[0049] In step S101, an RGB image sequence of a pedestrian is obtained.

[0050] The embodiments of the present invention are applicable to computing devices, such as personal computers, servers, etc. In the embodiments of the present invention, first, existing pedestrian videos can be collected. These videos can come from different channels, such as the storage devices of monitoring systems, the shooting records of mobile terminals, or public data sets, etc. Alternatively, pedestrian videos can be collected in real time through image acquisition devices. Then, according to a preset time interval (such as extracting one frame of image from the video every 0.5 seconds) or frame number interval (such as extracting one frame of image every 2 frames), images are extracted from the pedestrian videos to ensure that the extracted image sequence can accurately and coherently reflect the walking process and posture changes of the pedestrian. Finally, the extracted images are saved in RGB format to form an RGB image sequence of the pedestrian.

[0051] In step S102, the pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence.

[0052] In the embodiments of the present invention, the RGB image sequence is input into the pre-trained latent diffusion model (LatentDiffusion Models, LDMs). The pre-trained LDMs have learned certain feature representations and denoising capabilities on a large amount of data. Here, the knowledge of the LDMs is used to denoise the RGB image sequence once to obtain the diffusion features of each frame of RGB image in the RGB image sequence, so that the diffusion features have the overall shape information of the human body while naturally removing some color texture information irrelevant to the identity in the RGB image.

[0053] When denoising an RGB image sequence using a pre-trained latent diffusion model, preferably, first use the image encoder in the latent diffusion model to project each frame of the RGB image in the RGB image sequence into the latent space to obtain corresponding latent features; then, according to the denoising time step determined by training the latent diffusion model, use the U-net network in the latent diffusion model to perform one-step denoising on the latent features to obtain diffusion features.

[0054] In the embodiments of the present invention, the LDMs include an image encoder (Encoder) and a U-net network, and the U-net network is modified so that it runs once. Here, for the RGB image sequence , first use the image encoder to project each frame of the RGB image in the RGB image sequence into a low-dimensional latent space to compress the RGB image and preliminarily filter out high-frequency noise to obtain corresponding latent features, that is , then, perform diffusion in the latent space and predict the noise to be removed through the U-Net. Specifically, the U-Net directly simulates the denoising process at a given denoising time step stage through a single forward calculation to obtain the predicted result of the noise feature at this stage, that is, the diffusion feature. This diffusion feature mainly reflects the residual structure (i.e., the human shape) after the degradation of high-frequency details. The above process of denoising using the latent diffusion model is specifically expressed as , where represents the th frame of the RGB image 's diffusion feature, represents the image encoder, represents the U-Net, represents the denoising time step, R represents the set of real numbers, represents the spatial resolution and number of channels of the diffusion feature, , respectively represent the height and width of the RGB image, is the downsampling factor introduced by the image encoder , represents the total number of RGB images in the RGB image sequence.

[0055] In a feasible embodiment, before denoising the RGB image sequence using a pre-trained latent diffusion model, by training the latent diffusion model, the optimal denoising time step is determined so that the high-frequency details (such as color textures) and low-frequency structures (such as human contours) in the noise predicted by the LDMs reach a balance, that is, while suppressing the noise irrelevant to the identity, the complete and clear key overall shape information of the human body is retained.

[0056] In step S103, intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on each diffusion feature to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. Both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields.

[0057] In an embodiment of the present invention, an intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on each diffusion feature by using a feature matching module. The feature matching module is a learnable geometry-driven module, which aims to compress the multi-channel features of each pixel into a two-dimensional direction vector, so as to reduce the expression of RGB coding noise and highlight the local vectorization features of gait appearance and motion. Specifically, as Figure 2 shown, in the feature matching module, first, the features at the corresponding positions of the input features are matched, and then a direction assignment is performed to convert the features into two-dimensional direction vectors to enhance the local feature expression of the human body, obtaining a two-dimensional direction vector field. Among them, when the input feature is the diffusion feature corresponding to a single-frame image in the RGB image sequence, intra-frame feature matching is performed to generate a two-dimensional direction vector field representing the static features of the human body, that is, the static gait feature field. When the input feature is the diffusion feature corresponding to a single-frame image in the RGB image sequence and the diffusion features corresponding to the front and rear frame images adjacent to this frame image, inter-frame feature matching is performed to generate a two-dimensional direction vector field representing the dynamic features of the human body, that is, the dynamic gait feature field. The feature matching module executes intra-frame feature matching and inter-frame feature matching in parallel.

[0058] In a feasible embodiment, the intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on each diffusion feature through the following steps:

[0059] (1) According to the diffusion feature and the human silhouette pre-extracted from the RGB image corresponding to the diffusion feature, a query feature in the spatial dimension and a key feature in the time dimension of the RGB image corresponding to the diffusion feature are generated by using a preset feature generation strategy;

[0060] In an embodiment of the present invention, the feature matching module includes a first sub-module and a second sub-module. The first sub-module is used to generate query features in the spatial dimension of the RGB image corresponding to the diffusion feature, and the second sub-module is used to generate key features in the temporal dimension of the RGB image corresponding to the diffusion feature. The intra-frame feature matching and inter-frame feature matching follow a similar workflow. During the feature matching process, a human silhouette pre-extracted from the RGB image corresponding to the diffusion feature is introduced for constraint to cover the background area, thereby reducing background interference and allowing the gait feature field to represent only the features on the human body. Finally, the query features in the spatial dimension and the key features in the temporal dimension of the RGB image corresponding to the diffusion feature generated by the intra-frame feature matching mechanism, as well as the query features in the spatial dimension and the key features in the temporal dimension of the RGB image corresponding to the diffusion feature generated by the inter-frame feature matching mechanism are obtained.

[0061] In a feasible embodiment, the generation of query features and key features is achieved through the following steps:

[0062] (1.1) According to the diffusion feature and the human silhouette, use the query feature calculation formula to calculate the query features, where represents the query features corresponding to the th frame RGB image, represents the human silhouette of the th frame RGB image, represents the diffusion feature of the th frame RGB image, represents the first sub-module;

[0063] (1.2) According to the preset time reception range, use the key feature calculation formula to calculate the key features, where represents the key features corresponding to the th frame RGB image, represents the time reception range, represents the second sub-module.

[0064] In an embodiment of the present invention, the feature matching module includes a first sub-module and a second sub-module , and the first sub-module and the second sub-module are convolutional stacks with the same architecture but independently trained weights, that is, and will adjust their respective weight parameters according to their respective input data and target outputs during the training process to learn different feature representations. Finally, and It has its own unique weight parameters, and the two do not share weights, thus increasing the flexibility and expressive power of the model. Specifically, as Figure 3 shown and are both convolutional stacks stacked by four convolutional layers, and and have the number of input / output feature channels set to 4 and C respectively. Both are designed to extract and transform the features of their respective input data. Specifically, the data of the 4 input channels go through the layer-by-layer processing of the four convolutional layers in each convolutional stack. Each layer will perform a convolution operation on the input features through its independently trained weights to extract features at different levels. As the data is passed through the stack, the degree of abstraction of the features gradually increases, and finally, feature maps of C channels are output. These feature maps contain important feature information of the input data at different spatial / temporal positions and channel combinations. Based on and the output features, as shown in (a) of Figure 2 , the intra-frame feature matching and inter-frame feature matching follow a similar workflow as follows:

[0065] ,

[0066] ,

[0067] obtain the query features in the spatial dimension and the key features in the temporal dimension of the RGB image corresponding to the diffusion feature generated by using the intra-frame feature matching mechanism (i.e., ), and the query features in the spatial dimension and the key features in the temporal dimension of the RGB image corresponding to the diffusion feature generated by using the inter-frame feature matching mechanism (i.e., ). Among them, ), where corresponds to the RGB image, represents the query feature, represents the key feature, , , represents the human silhouette of the -th frame RGB image, and , represents the time reception range, represents the RGB image corresponding to the frame separated by frames from the -th frame in the temporal dimension, represents intra-frame feature matching, represents inter-frame feature matching. For the convenience of distinction and description, the query features and key features generated by using the intra-frame feature matching mechanism are respectively represented as , The query features and key features generated by the inter-frame feature matching mechanism are respectively denoted as and .

[0068] (2) According to the key features, direction assignment is performed on each pixel in the query features to obtain a static gait feature field and a dynamic gait feature field.

[0069] In the embodiment of the present invention, as shown in (b) of Figure 2 , through the Assignment of Direction (AoD) sub-module in the feature matching module, direction assignment is performed according to the query features and the key features to obtain a static gait feature field and a dynamic gait feature field. The static gait feature field is denoted as , and the dynamic gait feature field is denoted as , where represents the AoD sub-module.

[0070] In a feasible embodiment, the direction assignment of each pixel in the query features according to the key features is implemented through the following steps:

[0071] (2.1) Obtain the set of neighboring pixels of each pixel in the query features within the key features;

[0072] In the embodiment of the present invention, as shown in (b) of Figure 2 , for the sake of clarity, the frame index is ignored in the following description . Here, for a single pixel within the query features , denoted as , according to the preset height step and width step, its set of neighboring pixels within the key features is identified, denoted as , where represents the set of neighboring pixels, represents a single pixel within represents the height step, represents the width step. For simplicity, is regarded as a matrix, and its elements are arranged in raster order, that is, from top to bottom and from left to right. Therefore, has a shape of .

[0073] (2.2) Calculate the similarity distribution between each pixel in the query features and its corresponding set of neighboring pixels;

[0074] In the embodiments of the present invention, each pixel in the query feature is calculated for the similarity distribution between it and its neighboring pixels. Specifically, , where represents the similarity distribution, and , and the Softmax activation function operates along the last dimension.

[0075] (2.3) According to the similarity distribution and a preset direction template, each pixel in the query feature is assigned a direction.

[0076] In the embodiments of the present invention, in order to determine the final direction vector, a direction template pre-constructed according to a preset height step and width step is introduced. The direction template has an element correspondence relationship with the neighboring pixel set . Specifically, the direction template is represented as , and is regarded as a matrix with elements arranged in raster order, that is, from top to bottom and from left to right. Then has a shape of . Here, according to the similarity distribution and the direction template , a direction vector is assigned to each pixel in the query feature, that is , where represents the direction vector characterizing the human body feature. Finally, all the direction vectors obtained by performing intra-frame feature matching constitute the static gait feature field , and all the direction vectors obtained by performing inter-frame feature matching constitute the dynamic gait feature field . .

[0077] In a feasible embodiment, during the process of extracting the static gait feature field, the direction vectors with high response values are probabilistically set to zero through a random zero-padding strategy. Specifically, for the direction vectors in the static gait feature field whose vector modulus exceeds a preset threshold, these direction vectors with high response values are set to zero with a probability of 0.5, which is represented as , where represents the modulus of the direction vector , and represents the Bernoulli distribution, thereby suppressing noise, promoting textureless learning of gait features, enhancing the model's tolerance to noise, and improving the generalization performance.

[0078] In the above steps (1) and (2), the in-frame and inter-frame feature matching mechanisms of the geometric constraint-based feature matching module are used to assign directions to adjacent positions according to the feature similarities in the spatial and temporal dimensions respectively, so as to perform secondary denoising on the diffusion features predicted by the potential diffusion model, thereby effectively capturing the vectorized features of gait appearance and motion and improving the extraction quality of gait representation.

[0079] In step S104, the gait of the pedestrian is recognized based on the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

[0080] In the embodiment of the present invention, the dynamic gait feature field and the static gait feature field are input into the gait recognition network that adapts to dual-branch input and adopts a selective fusion attention mechanism in parallel. The gait of the pedestrian is recognized by this gait recognition network to obtain the gait features of the pedestrian. Specifically, the SoTA (State-of-the-Art, SoTA) model GaitBase in the current gait recognition field is used as the benchmark of the gait recognition network, which includes multi-stage feature extraction (Stage1 to Stage4) and a classification head (Gait Head). Here, GaitBase is modified to adapt to dual-branch input, and feature fusion is performed through a selective fusion attention mechanism after Stage3. Stage4 and Gait Head are the same as the benchmark GaitBase. The structure of the modified GaitBase is as Figure 4 shown.

[0081] In the embodiment of the present invention, gait recognition is regarded as a dynamic denoising process. Specifically, an RGB image sequence of the pedestrian is obtained, and the pre-trained potential diffusion model is used to perform denoising processing on the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on each diffusion feature to obtain the static gait feature field and the dynamic gait feature field of the pedestrian. The static gait feature field and the dynamic gait feature field are both two-dimensional direction vector fields. The gait of the pedestrian is recognized based on the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, cascade denoising is realized through the knowledge-driven potential diffusion model and the geometry-driven feature matching module, gradually stripping redundant information irrelevant to identity (including dynamic background interference, clothing texture changes, and local occlusions, etc.), and improving the extraction quality of gait representation and the accuracy of gait recognition.

[0082] Embodiment 2:

[0083] Figure 5 The structure of the gait recognition device provided in Embodiment 2 of the present invention is shown. For the sake of illustration, only the parts related to the embodiment of the present invention are shown, including:

[0084] An image acquisition unit 51 for acquiring an RGB image sequence of a pedestrian;

[0085] An image denoising unit 52 for denoising the RGB image sequence using a pre-trained latent diffusion model to obtain the diffusion features of each frame of RGB image in the RGB image sequence;

[0086] A feature matching unit 53 for performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each diffusion feature respectively to obtain a static gait feature field and a dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields;

[0087] A gait recognition unit 54 for recognizing the gait of the pedestrian based on the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

[0088] Preferably, the image denoising unit 52 includes:

[0089] An image projection unit for projecting each frame of RGB image in the RGB image sequence into the latent space using the image encoder in the latent diffusion model to obtain corresponding latent features;

[0090] A feature denoising unit for performing one-step denoising on the latent features using the U-net network in the latent diffusion model according to the denoising time step determined by training the latent diffusion model to obtain diffusion features.

[0091] The feature matching unit 53 includes:

[0092] A feature generation unit for generating a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature using a preset feature generation strategy according to the diffusion feature and the human silhouette pre-extracted from the RGB image corresponding to the diffusion feature;

[0093] A direction assignment unit for assigning directions to each pixel in the query feature according to the key feature to obtain a static gait feature field and a dynamic gait feature field.

[0094] In the embodiments of the present invention, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device is divided into different functional units or modules to implement all or part of the functions described above. Each unit and module of the device can be implemented by corresponding hardware or software units. Each unit and module can be an independent software or hardware unit, or can be integrated into a software or hardware unit, which is not used to limit the present invention. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the device can refer to the corresponding description in the foregoing method embodiments and will not be elaborated herein.

[0095] Embodiment Three:

[0096] Figure 6 The structure of the computing device provided in Embodiment Three of the present invention is shown. For the convenience of description, only the parts related to the embodiments of the present invention are shown.

[0097] The computing device 6 in the embodiments of the present invention includes a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, the steps in the foregoing method embodiment of a gait recognition method are implemented, for example Figure 1 Steps S101 to S104 shown. Alternatively, when the processor 60 executes the computer program 62, the functions of each unit in the foregoing device embodiments are implemented, for example Figure 5 The functions of the shown units.

[0098] In the embodiments of the present invention, an RGB image sequence of a pedestrian is acquired, and a pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. Both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, improving the extraction quality of gait representation and the accuracy of gait recognition.

[0099] The computing device in the embodiments of the present invention can be a personal computer or a server. The steps implemented when the processor 60 in the computing device 6 executes the computer program 62 to implement a gait recognition method can refer to the description in the foregoing method embodiments and will not be elaborated herein.

[0100] Embodiment 4:

[0101] In an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned embodiment of a gait recognition method are implemented. For example, Figure 1 the steps S101 to S104 shown. Alternatively, when the computer program is executed by a processor, the functions of each unit in the above-mentioned device embodiments are implemented. For example Figure 5 the functions of the units shown.

[0102] In an embodiment of the present invention, an RGB image sequence of a pedestrian is obtained, and a pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. Among them, both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, improving the extraction quality of gait representation and the accuracy of gait recognition.

[0103] The computer-readable storage medium in the embodiment of the present invention may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EEPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiment of the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0104] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the scope of disclosure involved in the above embodiments is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0105] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limitations on the scope of the invention. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.

Claims

1. A gait recognition method, characterized in that, The method includes the following steps: Obtain the RGB image sequence of a pedestrian; Use a pre-trained latent diffusion model to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence; Perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; Identify the gait of the pedestrian based on the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian; Among them, the step of performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian includes: According to the diffusion features and the human silhouette pre-extracted from the RGB image corresponding to the diffusion features, use a preset feature generation strategy to generate the query features in the spatial dimension and the key features in the time dimension of the RGB image corresponding to the diffusion features; According to the key features, assign directions to each pixel in the query features to obtain the static gait feature field and the dynamic gait feature field; The step of using a preset feature generation strategy to generate the query features in the spatial dimension and the key features in the time dimension of the RGB image corresponding to the diffusion features includes: According to the diffusion feature and the human silhouette, use the query feature calculation formula to calculate the query feature, where represents the query feature corresponding to the th frame of RGB image, represents the human silhouette of the th frame of RGB image, represents the diffusion feature of the th frame of RGB image, represents the first sub-module; Receive according to a preset time reception range, and use the key feature calculation formula to calculate the key feature, where represents the key feature corresponding to the nth frame of RGB image, represents the time reception range, represents the second sub-module.

2. The method according to claim 1, characterized in that, The step of using a pre-trained latent diffusion model to denoise the RGB image sequence includes: Use the image encoder in the latent diffusion model to project each frame of RGB image in the RGB image sequence into the latent space to obtain the corresponding latent features; According to the denoising time step determined by training the latent diffusion model, use the U-net network in the latent diffusion model to perform one-step denoising on the latent features to obtain the diffusion features.

3. The method according to claim 1, wherein The step of assigning directions to each pixel in the query features according to the key features includes: Obtain the set of neighboring pixels of each pixel in the query features within the key features; Calculate the similarity distribution between each pixel in the query features and its corresponding set of neighboring pixels; According to the similarity distribution and a preset direction template, assign directions to each pixel in the query features.

4. A gait recognition device, characterized in that, The device includes: An image acquisition unit for obtaining the RGB image sequence of a pedestrian; An image denoising unit for using a pre-trained latent diffusion model to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence; A feature matching unit for performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; A gait recognition unit, configured to recognize the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field, so as to obtain the gait features of the pedestrian; Wherein, the feature matching unit includes: A feature generation unit, configured to generate a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature by using a preset feature generation strategy according to the diffusion feature and a human silhouette pre-extracted from the RGB image corresponding to the diffusion feature, including: according to the diffusion feature and the human silhouette, using a query feature calculation formula to calculate the query feature, where represents the query feature corresponding to the th frame of RGB image, represents the human silhouette of the th frame of RGB image, represents the diffusion feature of the th frame of RGB image, represents the first sub-module; according to a preset time reception range, using a key feature calculation formula to calculate the key feature, where represents the key feature corresponding to the th frame of RGB image, represents the time reception range, represents the second sub-module; A direction assignment unit, configured to assign a direction to each pixel in the query feature according to the key feature, so as to obtain the static gait feature field and the dynamic gait feature field.

5. The device according to claim 4, characterized in that, The image denoising unit includes: An image projection unit, configured to project each frame of RGB image in the RGB image sequence into a latent space by using an image encoder in the latent diffusion model, so as to obtain corresponding latent features; A feature denoising unit, configured to perform one-step denoising on the latent features by using a U-net network in the latent diffusion model according to a denoising time step determined by training the latent diffusion model, so as to obtain the diffusion features.

6. A computing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Gait recognition method, device and equipment based on multiple modes and storage medium

    CN116721438A

  • System and method for in motion identification

    US20180232569A1