Gait recognition method, device and equipment and storage medium

By denoising the RGB image sequence using a potential diffusion model, and combining the feature matching technology of geometric constraints, static and dynamic gait feature fields are generated, which solves the problem of low recognition accuracy in the existing gait recognition methods, and achieves higher quality gait feature extraction and recognition.

CN120198970AActive Publication Date: 2025-06-24SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510669567.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing gait recognition methods have shortcomings in removing background interference and extracting identity characteristics, resulting in low recognition accuracy.

Method used

The pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain diffusion features, and then static and dynamic gait feature fields are generated through intra- and inter-frame feature matching of geometric constraints, and finally gait recognition is performed based on these feature fields.

Benefits of technology

By removing identity-independent noise information, the extraction quality and recognition accuracy of gait features are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198970A_ABST
    Figure CN120198970A_ABST
Patent Text Reader

Abstract

The invention is applicable to the technical field of biological feature recognition, and provides a gait recognition method, device and equipment and a storage medium, and the method comprises the steps: obtaining an RGB image sequence of a pedestrian, carrying out the denoising of the RGB image sequence through employing a pre-trained potential diffusion model, obtaining the diffusion feature of each frame of RGB image in the RGB image sequence, and obtaining a recognition result of the RGB image; each diffusion feature is subjected to geometrically constrained intra-frame feature matching and inter-frame feature matching, a static gait feature field and a dynamic gait feature field of the pedestrian are obtained, the static gait feature field and the dynamic gait feature field are both two-dimensional direction vector fields, and the static gait feature field and the dynamic gait feature field are used as direction vector fields; the gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field, the gait features of the pedestrian are obtained, and therefore gait feature extraction is achieved through the denoising process of continuously removing features irrelevant to the identity, and the extraction quality of gait representation and the accuracy of gait recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biometric recognition, and particularly relates to a gait recognition method, device, equipment and storage medium. Background Art

[0002] The differences among pedestrians in videos are very subtle because the videos are in a very high-dimensional space and there is a lot of noise unrelated to identity, which makes it a very challenging task to extract corresponding biometric features (gait recognition) from pedestrian videos and has attracted increasing attention.

[0003] Current gait recognition methods can be mainly divided into three technical routes: unimodal gait recognition, multimodal gait recognition, and RGB-based gait recognition. Among them, unimodal gait recognition methods usually first use pre-designed networks (such as human segmentation networks, human semantic analysis networks, human key point prediction networks, etc.) to extract gait modalities related to identity, remove noise information unrelated to identity such as background and clothing color through this explicit preprocessing, and then input the obtained gait modalities (such as silhouette maps, pose heat maps, etc.) into the gait recognition network for corresponding recognition. Although this separated design can effectively reduce background interference, it may cause loss of identity-related information; multimodal gait recognition methods improve the recognition accuracy by integrating multiple complementary gait modalities (such as combining contour maps, skeleton sequences, and semantic segmentation maps). Although the fusion of multi-source information can theoretically enhance the feature expression ability, these methods rely on third-party pre-trained models to obtain modal features. Since these models are limited by original supervision tasks unrelated to identity (such as pose estimation, semantic segmentation), the extracted modal features still have the inherent defect of insufficient correlation with identity, and there is a lack of a unified feature space alignment mechanism, resulting in a bottleneck in improving the accuracy of multimodal fusion; for RGB-based gait recognition, there are now many RGB-based end-to-end gait recognition methods. These methods usually draw on existing gait modal features (such as silhouettes, SMPL models, human skeleton points, etc.), directly extract features from RGB images, and input them into the downstream gait recognition network for recognition. Although these methods can retain richer visual information, they require complex designs to simulate the functions of professional preprocessing networks and fail to effectively establish a direct mapping relationship between image pixel-level information and identity features, resulting in difficulty in effectively separating identity features from interference factors. Therefore, there is an urgent need for a new gait recognition method to solve the above problems. Summary of the Invention

[0004] The purpose of the present invention is to provide a gait recognition method, device, equipment and storage medium, aiming to solve the problem of low gait recognition accuracy caused by the existing technology.

[0005] In a first aspect, the present invention provides a gait recognition method, and the method includes the following steps: Obtain an RGB image sequence of a pedestrian; Use a pre-trained latent diffusion model to denoise the RGB image sequence to obtain diffusion features of each frame of RGB image in the RGB image sequence; Perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; Recognize the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait feature of the pedestrian.

[0006] In some embodiments, the step of using a pre-trained latent diffusion model to denoise the RGB image sequence includes: Use the image encoder in the latent diffusion model to project each frame of RGB image in the RGB image sequence into the latent space to obtain corresponding latent features; According to the denoising time step determined by training the latent diffusion model, use the U-net network in the latent diffusion model to perform one-step denoising on the latent features to obtain the diffusion features.

[0007] In some embodiments, the step of performing intra-frame feature matching and inter-frame feature matching with frame geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian includes: According to the diffusion features and the human silhouette pre-extracted from the RGB image corresponding to the diffusion features, use a preset feature generation strategy to generate query features in the spatial dimension and key features in the time dimension of the RGB image corresponding to the diffusion features; According to the key features, assign directions to each pixel in the query features to obtain the static gait feature field and the dynamic gait feature field.

[0008] In some embodiments, the step of using a preset feature generation strategy to generate query features in the spatial dimension and key features in the time dimension of the RGB image corresponding to the diffusion features includes: According to the diffusion features and the human silhouette, use the query feature calculation formula to calculate the query features, where represents the query feature corresponding to the th frame of RGB image, represents the human silhouette of the th frame of RGB image, Represents the diffusion feature of the RGB image of the frame, indicating the first sub-module; According to the preset time reception range, using the key feature calculation formula to calculate the key feature, where represents the key feature corresponding to the RGB image of the frame, and represents the time reception range, indicating the second sub-module.

[0009] In some embodiments, the step of performing direction assignment on each pixel in the query feature according to the key feature includes: Obtaining the set of neighboring pixels of each pixel in the query feature within the key feature; Calculating the similarity distribution between each pixel in the query feature and its corresponding set of neighboring pixels; Performing direction assignment on each pixel in the query feature according to the similarity distribution and a preset direction template.

[0010] In a second aspect, the present invention provides a gait recognition device, which includes: An image acquisition unit for acquiring an RGB image sequence of a pedestrian; An image denoising unit for denoising the RGB image sequence by using a pre-trained latent diffusion model to obtain the diffusion feature of each frame of RGB image in the RGB image sequence; A feature matching unit for performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; A gait recognition unit for recognizing the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait feature of the pedestrian.

[0011] In some embodiments, the image denoising unit includes: An image projection unit for projecting each frame of RGB image in the RGB image sequence into the latent space by using the image encoder in the latent diffusion model to obtain the corresponding latent feature; A feature denoising unit for performing one-step denoising on the latent feature by using the U-net network in the latent diffusion model according to the denoising time step determined by training the latent diffusion model to obtain the diffusion feature.

[0012] In some embodiments, the feature matching unit includes: A feature generation unit, configured to generate a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature by using a preset feature generation strategy according to the diffusion feature and a human silhouette pre-extracted from the RGB image corresponding to the diffusion feature; A direction assignment unit, configured to perform direction assignment on each pixel in the query feature according to the key feature to obtain the static gait feature field and the dynamic gait feature field.

[0013] In a third aspect, the present invention further provides a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0015] When performing gait recognition in an embodiment of the present invention, an RGB image sequence of a pedestrian is acquired, and a pre-trained latent diffusion model is used to perform denoising processing on the RGB image sequence to obtain diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. The static gait feature field and the dynamic gait feature field are both two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait feature of the pedestrian. Thus, gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, improving the extraction quality of gait representation and the accuracy of gait recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 is a schematic flowchart of a gait recognition method provided in Embodiment 1 of the present invention; Figure 2 is a schematic structural diagram of a feature matching module in a gait recognition method provided in Embodiment 1 of the present invention; Figure 3 is a schematic structural diagram of the first / second sub-module in the feature matching module provided in Embodiment 1 of the present invention; Figure 4 is a schematic structural diagram of an improved GaitBase provided in Embodiment 1 of the present invention; Figure 5 is a schematic structural diagram of a gait recognition device provided in Embodiment 2 of the present invention; Figure 6 It is a schematic structural diagram of a computing device provided in Embodiment 3 of the present invention. Detailed implementation manners

[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their combinations. And the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms. Terms such as "first", "second" and similar words do not denote any order, quantity or importance, but are only used to distinguish different components. "Connection" or "coupling" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. The term "plurality" means two or more, and other quantifiers are similar thereto.

[0019] In order to keep the following description of the embodiments of the present invention clear and concise, the detailed descriptions of some known functions and known components are omitted in this specification.

[0020] The following describes in detail the specific implementation of the present invention with reference to specific embodiments: Embodiment 1: Figure 1 The implementation process of the gait recognition method provided in Embodiment 1 of the present invention is shown. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown and are described in detail as follows: In step S101, an RGB image sequence of a pedestrian is acquired.

[0021] Embodiments of the present invention are applicable to computing devices, such as personal computers, servers, etc. In embodiments of the present invention, first, existing pedestrian videos can be collected. These videos can come from different channels, such as the storage devices of surveillance systems, the shooting records of mobile terminals, or public data sets, etc. Alternatively, pedestrian videos can be collected in real time through image acquisition devices. Then, according to a preset time interval (such as extracting one frame of image from the video every 0.5 seconds) or frame number interval (such as extracting one frame every 2 frames), images are extracted from the pedestrian videos to ensure that the extracted image sequence can accurately and coherently reflect the walking process and posture changes of pedestrians. Finally, the extracted images are saved in RGB format to form an RGB image sequence of pedestrians.

[0022] In step S102, a pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence.

[0023] In embodiments of the present invention, the RGB image sequence is input into a pre-trained latent diffusion model (LatentDiffusion Models, LDMs). The pre-trained LDMs have learned certain feature representations and denoising capabilities on a large amount of data. Here, the knowledge of the LDMs is used to denoise the RGB image sequence once to obtain the diffusion features of each frame of RGB image in the RGB image sequence, so that the diffusion features have the overall shape information of the human body while naturally removing some color texture information irrelevant to identity in the RGB image.

[0024] When using a pre-trained latent diffusion model to denoise the RGB image sequence, preferably, first, the image encoder in the latent diffusion model is used to project each frame of RGB image in the RGB image sequence into the latent space to obtain the corresponding latent features; then, according to the denoising time step determined by training the latent diffusion model, the U-net network in the latent diffusion model is used to denoise the latent features once to obtain the diffusion features.

[0025] In embodiments of the present invention, the LDMs include an image encoder (Encoder) and a U-net network, and the U-net network is modified so that it runs once. Here, for the RGB image sequence , first, the image encoder is used to project each frame of RGB image in the RGB image sequence into a low-dimensional latent space to compress the RGB image and preliminarily filter high-frequency noise to obtain the corresponding latent features, that is , then, in the latent space Diffusion is performed, and the noise to be removed is predicted by U-Net. Specifically, U-Net directly simulates the denoising process at a given denoising time step through a single forward calculation, obtaining the predicted result of the noise characteristics at this stage, that is, the diffusion characteristics. The diffusion characteristics mainly reflect the residual structure (i.e., the human shape) after the degradation of high-frequency details. The above process of denoising using the latent diffusion model is specifically expressed as , where represents the diffusion characteristics of the th frame RGB image represents the image encoder, represents U-Net, represents the denoising time step, R represents the set of real numbers, represents the spatial resolution and number of channels of the diffusion characteristics, , respectively represent the height and width of the RGB image, is the downsampling factor introduced by the image encoder represents the total number of frames of RGB images in the RGB image sequence.

[0026] In a feasible embodiment, before using the pre-trained latent diffusion model to denoise the RGB image sequence, the optimal denoising time step is determined by training the latent diffusion model, so that the high-frequency details (such as color texture) and low-frequency structures (such as human contours) in the noise predicted by LDMs reach a balance, that is, while suppressing the noise irrelevant to identity, the complete and clear overall shape information of the human body key is retained.

[0027] In step S103, intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on each diffusion feature to obtain the static gait feature field and dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields.

[0028] In the embodiment of the present invention, a feature matching module is used to perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each diffusion feature respectively. The feature matching module is a learnable geometric driving module, aiming to compress the multi-channel features of each pixel into a two-dimensional direction vector, thereby reducing the expression of RGB coding noise and highlighting the local vectorized features of gait appearance and movement. Specifically, as Figure 2As shown in the figure, in the feature matching module, first, the features at the corresponding positions of the input features are matched, and then an assignment in one direction is performed to convert the features into two-dimensional direction vectors to enhance the local feature expression of the human body, obtaining a two-dimensional direction vector field. Among them, when the input feature is the diffusion feature corresponding to a single-frame image in the RGB image sequence, in-frame feature matching is performed to generate a two-dimensional direction vector field representing the static features of the human body, that is, the static gait feature field. When the input feature is the diffusion feature corresponding to a single-frame image in the RGB image sequence and the diffusion features corresponding to the front and rear frame images adjacent to this frame image, inter-frame feature matching is performed to generate a two-dimensional direction vector field representing the dynamic features of the human body, that is, the dynamic gait feature field. The feature matching module performs in-frame feature matching and inter-frame feature matching in parallel.

[0029] In a feasible embodiment, the in-frame feature matching and inter-frame feature matching with geometric constraints on each diffusion feature are implemented through the following steps: (1) According to the diffusion feature and the human silhouette pre-extracted from the RGB image corresponding to the diffusion feature, a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature are generated by using a preset feature generation strategy. In the embodiment of the present invention, the feature matching module includes a first sub-module and a second sub-module. The first sub-module is used to generate a query feature in the spatial dimension of the RGB image corresponding to the diffusion feature, and the second sub-module is used to generate a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature. The in-frame feature matching and inter-frame feature matching follow a similar working process. During the feature matching process, the human silhouette pre-extracted from the RGB image corresponding to the diffusion feature is introduced for constraint to cover the background area, thereby reducing background interference and making the gait feature field only represent the features on the human body. Finally, the query feature in the spatial dimension and the key feature in the temporal dimension of the RGB image corresponding to the diffusion feature generated by using the in-frame feature matching mechanism, and the query feature in the spatial dimension and the key feature in the temporal dimension of the RGB image corresponding to the diffusion feature generated by using the inter-frame feature matching mechanism are obtained.

[0030] In a feasible embodiment, the generation of the query feature and the key feature is implemented through the following steps: (1.1) According to the diffusion feature and the human silhouette, use the query feature calculation formula to calculate the query feature, where represents the query feature corresponding to the th frame RGB image, represents the human silhouette of the th frame RGB image, represents the diffusion feature of the th frame RGB image, Indicates the first sub-module; (1.2) According to the preset time reception range, use the key feature calculation formula to calculate the key features, where represents the key features corresponding to the th frame of RGB image, represents the time reception range,

[0031] In the embodiments of the present invention, the feature matching module includes a first sub-module and a second sub-module , the first sub-module and the second sub-module are convolutional stacks (ConvolutionalStack) with the same architecture but independently trained weights, that is, and will adjust their respective weight parameters according to their respective input data and target outputs during the training process to learn different feature representations. Finally, and have their own unique weight parameters, and the two do not share weights, thus increasing the flexibility and expressive ability of the model. Specifically, as Figure 3 shown, and are both convolutional stacks composed of four stacked convolutional layers, and and set the input / output feature channels to 4 and C respectively. The two are designed to extract and transform the features of their respective input data. Specifically, the input 4-channel data is processed layer by layer through the four convolutional layers in each convolutional stack. Each layer will perform a convolutional operation on the input features through its independently trained weights to extract different levels of features. As the data is passed through the stack, the degree of abstraction of the features gradually increases, and finally C-channel feature maps are output. These feature maps contain important feature information of the input data at different spatial / temporal positions and channel combinations. Based on and output features, as Figure 2 shown in (a), the intra-frame feature matching and inter-frame feature matching follow a similar workflow as follows: , , obtain the query features in the spatial dimension and the key features in the temporal dimension of the RGB image corresponding to the diffusion feature generated by using the intra-frame feature matching mechanism (i.e., ) and , and the query features in the spatial dimension of the RGB image generated by adopting an inter-frame feature matching mechanism (i.e., ) and the key features in the temporal dimension corresponding to the diffusion features , where represents the query feature and represents the key feature. Among them, , , represents the human silhouette of the -th frame RGB image, and , represents the time reception range, represents the RGB image corresponding to the -th frame separated by frames in the temporal dimension, represents intra-frame feature matching, represents inter-frame feature matching. For the convenience of distinction and description, the query features and key features generated by adopting the intra-frame feature matching mechanism are respectively represented as , , and the query features and key features generated by adopting the inter-frame feature matching mechanism are respectively represented as , .

[0032] (2) According to the key features, direction assignment is performed on each pixel in the query features to obtain a static gait feature field and a dynamic gait feature field.

[0033] In the embodiment of the present invention, as shown in (b) of Figure 2 , through the Assignment of Direction (AoD) sub-module in the feature matching module, direction assignment is performed according to the query features and the key features to obtain a static gait feature field and a dynamic gait feature field. The static gait feature field is represented as , and the dynamic gait feature field is represented as , represents the AoD sub-module.

[0034] In a feasible embodiment, the direction assignment of each pixel in the query features according to the key features is realized through the following steps: (2.1) Obtain the set of neighboring pixels of each pixel in the query features within the key features; In the embodiment of the present invention, as shown in (b) of Figure 2 , for the sake of clarity, the frame index is ignored in the following description. Here, for a single pixel within the query features ​ , denoted as , identify its neighboring pixel set within the key feature , denoted as , where represents the neighboring pixel set, represents a single pixel within represents the height step, represents the width step. For simplicity, is regarded as a matrix, and its elements are arranged in raster order, that is, from top to bottom and from left to right. Therefore, has a shape of .

[0035] (2.2) Calculate the similarity distribution between each pixel in the query feature and its corresponding neighboring pixel set; In the embodiment of the present invention, calculate each pixel in the query feature and its similarity distribution with its neighboring pixels. Specifically, , where represents the similarity distribution, and , the Softmax activation function operates along the last dimension.

[0036] (2.3) Assign directions to each pixel in the query feature according to the similarity distribution and the preset direction template.

[0037] In the embodiment of the present invention, in order to determine the final direction vector, a direction template pre-constructed according to the preset height step and width step is introduced. This direction template has an element correspondence relationship with the neighboring pixel set . Specifically, the direction template is denoted as . Regarding as a matrix with elements arranged in raster order, that is, from top to bottom and from left to right, then has a shape of . Here, according to the similarity distribution and the direction template , assign a direction vector to each pixel in the query feature, that is, , where represents the direction vector characterizing the human body feature. Finally, all the direction vectors obtained by performing intra-frame feature matching constitute the static gait feature field , all the direction vectors obtained by performing inter-frame feature matching constitute a dynamic gait feature field .

[0038] In a feasible embodiment, during the process of extracting the static gait feature field, the direction vectors with high response values are probabilistically set to zero through a random zero-padding strategy. Specifically, for the direction vectors in the static gait feature field whose vector norm exceeds a preset threshold, these direction vectors with high response values are set to zero with a probability of 0.5, denoted as , where represents the vector norm of the direction vector , represents the Bernoulli distribution, thereby suppressing noise, promoting textureless learning of gait features, enhancing the model's tolerance to noise, and improving generalization performance.

[0039] The above steps (1) and (2) respectively assign directions to adjacent positions according to the feature similarity in the spatial and temporal dimensions through the intra-frame and inter-frame feature matching mechanisms of the geometric constraint-based feature matching module, so as to perform secondary denoising on the diffusion features predicted by the potential diffusion model, thereby effectively capturing the vectorized features of gait appearance and motion and improving the extraction quality of gait representation.

[0040] In step S104, the gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

[0041] In the embodiment of the present invention, the dynamic gait feature field and the static gait feature field are input in parallel into a gait recognition network that adapts to dual-branch input and adopts a selective fusion attention mechanism. The gait of the pedestrian is recognized by this gait recognition network to obtain the gait features of the pedestrian. Specifically, the current SoTA (State-of-the-Art, SoTA) model GaitBase in the field of gait recognition is used as the benchmark of the gait recognition network. It includes multi-stage feature extraction (Stage1 to Stage4) and a classification head (Gait Head). Here, GaitBase is modified to adapt to dual-branch input, and feature fusion is performed through a selective fusion attention mechanism after Stage3. Stage4 and Gait Head are the same as the benchmark GaitBase. The structure of the modified GaitBase is as Figure 4 shown.

[0042] In an embodiment of the present invention, gait recognition is regarded as a dynamic denoising process. Specifically, an RGB image sequence of a pedestrian is obtained, and the pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian. Both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, cascaded denoising is realized through the knowledge-driven latent diffusion model and the geometric-driven feature matching module, and redundant information irrelevant to identity (including dynamic background interference, clothing texture changes, and partial occlusions, etc.) is gradually stripped, improving the extraction quality of gait representation and the accuracy of gait recognition.

[0043] Embodiment 2: Figure 5 The structure of the gait recognition device provided in Embodiment 2 of the present invention is shown. For the convenience of description, only the parts related to Embodiment 2 of the present invention are shown, including: An image acquisition unit 51, configured to acquire an RGB image sequence of a pedestrian; An image denoising unit 52, configured to denoise the RGB image sequence by using the pre-trained latent diffusion model to obtain the diffusion features of each frame of RGB image in the RGB image sequence; A feature matching unit 53, configured to perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each diffusion feature respectively to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; A gait recognition unit 54, configured to recognize the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

[0044] Preferably, the image denoising unit 52 includes: An image projection unit, configured to project each frame of RGB image in the RGB image sequence into the latent space by using the image encoder in the latent diffusion model to obtain the corresponding latent features; A feature denoising unit, configured to perform one-step denoising on the latent features by using the U-net network in the latent diffusion model according to the denoising time step determined by training the latent diffusion model to obtain the diffusion features.

[0045] The feature matching unit 53 includes: A feature generation unit, configured to generate a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion feature by using a preset feature generation strategy according to the diffusion feature and a human silhouette pre-extracted from the RGB image corresponding to the diffusion feature; A direction assignment unit, configured to assign a direction to each pixel in the query feature according to the key feature, to obtain a static gait feature field and a dynamic gait feature field.

[0046] In the embodiments of the present invention, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used for illustration. In practical applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to implement all or part of the functions described above. Each unit and module of the device can be implemented by corresponding hardware or software units. Each unit and module can be an independent software or hardware unit, or integrated into a software or hardware unit, which is not used to limit the present invention. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the device can refer to the corresponding descriptions in the foregoing method embodiments and will not be elaborated herein.

[0047] Embodiment 3: Figure 6 The structure of a computing device provided in Embodiment 3 of the present invention is shown. For the convenience of description, only the parts related to Embodiment 3 of the present invention are shown.

[0048] The computing device 6 in the embodiments of the present invention includes a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, the steps in the foregoing method embodiment of a gait recognition method are implemented, such as Figure 1 the steps S101 to S104 shown. Alternatively, when the processor 60 executes the computer program 62, the functions of each unit in the foregoing device embodiments are implemented, such as Figure 5 the functions of the units shown.

[0049] In an embodiment of the present invention, an RGB image sequence of a pedestrian is obtained, and a pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. The static gait feature field and the dynamic gait feature field are both two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, improving the extraction quality of gait representation and the accuracy of gait recognition.

[0050] The computing device in the embodiment of the present invention may be a personal computer or a server. When the processor 60 in the computing device 6 executes the computer program 62 to implement a gait recognition method, the steps implemented may refer to the description of the foregoing method embodiment and will not be elaborated herein.

[0051] Embodiment 4: In an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the foregoing embodiment of a gait recognition method are implemented. For example, Figure 1 the steps S101 to S104 shown. Alternatively, when the computer program is executed by a processor, the functions of each unit in the foregoing device embodiments are implemented. For example, Figure 5 the functions of the units shown.

[0052] In an embodiment of the present invention, an RGB image sequence of a pedestrian is obtained, and a pre-trained latent diffusion model is used to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence. Intra-frame feature matching and inter-frame feature matching with geometric constraints are respectively performed on the diffusion features to obtain a static gait feature field and a dynamic gait feature field of the pedestrian. The static gait feature field and the dynamic gait feature field are both two-dimensional direction vector fields. The gait of the pedestrian is recognized according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian. Thus, gait feature extraction is realized through a denoising process of continuously removing identity-irrelevant features, improving the extraction quality of gait representation and the accuracy of gait recognition.

[0053] The computer-readable storage medium in the embodiments of the present invention may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EEPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0054] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the scope of disclosure involved in the above embodiments is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0055] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present invention. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

Claims

1. A gait recognition method, characterized in that, The method includes the following steps: Obtain an RGB image sequence of a pedestrian; Use a pre-trained latent diffusion model to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence; Perform intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; Identify the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

2. The method according to claim 1, characterized in that The step of using a pre-trained latent diffusion model to denoise the RGB image sequence includes: Use the image encoder in the latent diffusion model to project each frame of RGB image in the RGB image sequence into the latent space to obtain corresponding latent features; According to the denoising time step determined by training the latent diffusion model, use the U-net network in the latent diffusion model to perform one-step denoising on the latent features to obtain the diffusion features.

3. The method according to claim 1, wherein The step of performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian includes: According to the diffusion features and the human silhouette pre-extracted from the RGB image corresponding to the diffusion features, use a preset feature generation strategy to generate the query features in the spatial dimension and the key features in the time dimension of the RGB image corresponding to the diffusion features; According to the key features, assign directions to each pixel in the query features to obtain the static gait feature field and the dynamic gait feature field.

4. The method according to claim 3, characterized in that, The step of using a preset feature generation strategy to generate the query features in the spatial dimension and the key features in the time dimension of the RGB image corresponding to the diffusion features includes: According to the diffusion feature and the human silhouette, use the query feature calculation formula to calculate the query feature, where represents the query feature corresponding to the th frame of RGB image, represents the human silhouette of the th frame of RGB image, represents the diffusion feature of the th frame of RGB image, represents the first sub-module; Receive according to a preset time receiving range and calculate the key feature by using the key feature calculation formula to calculate the key feature, where represents the key feature corresponding to the nth frame of RGB image represents the time receiving range represents the second sub-module 5. The method according to claim 3, characterized in that, The step of assigning directions to each pixel in the query features according to the key features includes: Obtain the set of neighboring pixels of each pixel in the query features within the key features; Calculate the similarity distribution between each pixel in the query features and its corresponding set of neighboring pixels; Assign directions to each pixel in the query features according to the similarity distribution and a preset direction template.

6. A gait recognition device, characterized in that, The device includes: An image acquisition unit for obtaining an RGB image sequence of a pedestrian; An image denoising unit for using a pre-trained latent diffusion model to denoise the RGB image sequence to obtain the diffusion features of each frame of RGB image in the RGB image sequence; A feature matching unit for performing intra-frame feature matching and inter-frame feature matching with geometric constraints on each of the diffusion features to obtain the static gait feature field and the dynamic gait feature field of the pedestrian, where both the static gait feature field and the dynamic gait feature field are two-dimensional direction vector fields; A gait recognition unit for identifying the gait of the pedestrian according to the dynamic gait feature field and the static gait feature field to obtain the gait features of the pedestrian.

7. The device according to claim 6, characterized in that, The image denoising unit includes: An image projection unit, configured to project each frame of RGB image in the RGB image sequence into a latent space by using the image encoder in the latent diffusion model to obtain corresponding latent features; A feature denoising unit, configured to perform one-step denoising on the latent features by using the U-net network in the latent diffusion model according to the denoising time step determined by training the latent diffusion model to obtain the diffusion features.

8. The device according to claim 6, wherein, The feature matching unit includes: A feature generation unit, configured to generate a query feature in the spatial dimension and a key feature in the temporal dimension of the RGB image corresponding to the diffusion features according to the diffusion features and the human silhouette pre-extracted from the RGB image corresponding to the diffusion features by using a preset feature generation strategy; A direction assignment unit, configured to assign a direction to each pixel in the query feature according to the key feature to obtain the static gait feature field and the dynamic gait feature field.

9. A computing device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Gait recognition method, device and equipment based on multiple modes and storage medium

    CN116721438A

  • Gait recognition methods and systems

    US20140270402A1

  • System and method for in motion identification

    US20180232569A1

  • Method, apparatus, and system for human recognition based on gait features

    US20220026530A1

  • System and method for characterizing and monitoring health of an animal based on gait and postural movements

    US20230083421A1