Gait recognition method based on multi-modal feature fusion in low-illumination scene

By extracting gait profile sequence and skeleton features in low-light scenarios and performing multimodal feature fusion, the problem of low gait recognition accuracy under low light is solved, and higher recognition accuracy and robustness are achieved.

CN119942590AActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510006981.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

In low-light scenarios, existing gait recognition technology is difficult to extract the complete human contour, resulting in a greatly reduced recognition accuracy.

Method used

The gait recognition method based on multimodal feature fusion in low-light scenarios is adopted. By extracting the gait profile sequence and skeleton features from single-person walking videos and inputting them into the multimodal feature fusion network for feature fusion, the fused pedestrian gait features are generated to achieve recognition.

Benefits of technology

Under low light conditions, the accuracy, robustness and practicality of gait recognition are improved, allowing pedestrian identity to be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942590A_ABST
    Figure CN119942590A_ABST
Patent Text Reader

Abstract

The invention discloses a gait recognition method based on multi-modal feature fusion in a low-illumination scene, and belongs to the field of gait recognition. The method comprises the following steps: firstly, acquiring a batch of pictures from a collected video to generate an initial detection set, inputting the initial detection set into a figure detection extraction network to extract a gait profile diagram in the pictures, and carrying out noise reduction processing on the gait profile diagram to improve the definition of the gait profile diagram in a low-light scene; and inputting the initial detection set into a character skeleton feature extraction network to generate a skeleton model of pedestrians in the picture, inputting the processed gait contour map and the skeleton model into a multi-modal feature fusion network, and fusing two logic outputs in the fusion network to obtain a final pedestrian recognition result. According to the method, various kinds of information existing in the picture can be fully utilized under the low-illumination condition, comprehensive feature representation is obtained, and the accuracy, robustness and practicability of gait recognition under the low-illumination scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gait recognition, and in particular to a gait recognition method based on multimodal feature fusion in low-light scenarios. Background Art

[0002] Gait recognition is a biometric application that aims to identify pedestrians by their walking patterns. The field of deep learning regards gait recognition as a vision-based person retrieval method, that is, identifying a moving subject from a given gait sequence captured by a visual camera. Compared with other forms of biometric technology (such as facial recognition, fingerprint recognition, etc.), gait recognition has many significant advantages. The outstanding advantage of gait as a biometric feature is that it can be used for human recognition at a distance. In other words, gait can be used at low resolution. It usually occupies more pixels than other biometric features (such as face, fingerprint, iris). Gait information, as an inherent biometric feature, has unique advantages in gait recognition. By extracting key information from the pedestrian's gait sequence, such as the gait profile and the skeleton features of the gait, we can more accurately identify the identity of the pedestrian, thereby realizing the recognition of pedestrians at a distance by gait.

[0003] Although gait recognition has made significant progress in the past few years, there are still many challenges in practical applications. The main challenge is that the illumination of the recognition scene is too low. In low-light conditions, it is difficult to extract the complete human outline from the surveillance video obtained in real life. The low-quality outline has a great impact on the production of the template, which will make pedestrian recognition in low-light scenes extremely difficult, thereby greatly reducing the recognition accuracy.

[0004] The current gait recognition technology is mainly divided into two types: model-based and template-based. The gait energy image, which is the average value of continuous contours in a complete gait cycle, is the most commonly used template-based method, which requires image preprocessing from gait images or videos to distinguish human contours. The advantage of this method is that it can achieve reliable recognition results in a controllable environment, because these template features provide a lot of discriminative information for gait recognition. However, due to occlusion and illumination changes, it is difficult to obtain a complete human contour in real life. In the second category of model-based methods, a predefined human model is required to describe the dynamic and static characteristics of pedestrians. The main disadvantage of this category is that most methods only use sparse key point information, and there are still great challenges in extracting basic models from gait images and videos. Due to the lack of flexibility in fixed recognition methods and feature aggregation methods, these methods have obvious deficiencies in extracting gait contours and skeleton features in low-light conditions, and are difficult to apply to the recognition of people in low-light scenes. Summary of the invention

[0005] The present invention provides a gait recognition method based on multimodal feature fusion in low-light scenes, which can make full use of various information in the picture under low-light conditions, obtain a more comprehensive feature representation, and effectively improve the accuracy, robustness and practicality of gait recognition in low-light scenes.

[0006] An embodiment of the present invention provides a gait recognition method based on multimodal feature fusion in a low-light scene, comprising the following steps:

[0007] Step 1: obtain a gait contour sequence based on a single person walking video in a low-light scene, remodel the contour sequence based on a partial contour sequence and perform noise reduction, and input the denoised gait contour into a person detection and extraction network to obtain the pedestrian's gait contour features;

[0008] Step 2, inputting the gait profile sequence into a human skeleton feature extraction network to extract skeleton features of pedestrians in the image;

[0009] Step 3: Input the pedestrian's gait profile features and the skeleton features into a multimodal feature fusion network for feature fusion to obtain fused pedestrian gait features, and perform gait recognition based on the fused pedestrian gait features.

[0010] Optionally, in one embodiment of the present invention, step 1 further comprises:

[0011] Step 11, processing the CASIA-B dataset to simulate the gait of pedestrians in low-light scenes, and obtaining a single-person walking video in low-light scenes;

[0012] Step 12, using a pedestrian detection and segmentation algorithm to extract and segment the single-person walking image in the single-person walking video into a gait profile sequence;

[0013] Step 13, for the gait profile sequence, by synthesizing gait from partial sequence frames, the characteristics of the gait cycle are modeled as a Gaussian distribution, and the characteristics of the gait cycle are expressed as a continuous function, and a complete cycle gait is generated through modeling;

[0014] Step 14, denoising the reconstructed gait profile image, and inputting the denoised gait profile image into a person detection and extraction network to obtain the gait profile features of the pedestrian.

[0015] Optionally, in one embodiment of the present invention, step 13 further includes:

[0016] The dimension of the gait profile data is reduced and features are extracted. The formula is as follows:

[0017] ξλξ T =(SS T )×1 / N

[0018] Among them, ξ represents the common feature vector basis generated by the gait dataset, S represents the shape vector matrix, and N represents the total number of contours in the gait dataset;

[0019] The shape vector u i Project it onto the common eigenvalue basis vector ξ and obtain the corresponding eigenvalue a i , the formula is as follows:

[0020] a i =ξ T u i

[0021] Among them, ξ T represents the transpose of the common feature vector basis generated by the gait dataset, u i represents the shape vector;

[0022] The feature vector a i Transform it into a continuous function, and align the transformed eigenvectors and enhance them into a zero-mean eigenmatrix F, as follows:

[0023] F≡[a1(t),…,a M (t)] m×∞

[0024] Among them, a i (t) represents the feature vector a i The transformed eigenvector, m represents the number of eigenvectors, and ∞ represents that F has infinite dimensions;

[0025] The infinite-dimensional F matrix is ​​converted into a finite-dimensional feature covariance matrix as follows:

[0026] K M×M =F T F

[0027] Among them, K M×M represents a finite-dimensional eigenvector matrix, and F is a characteristic matrix with zero mean;

[0028] Calculate the components of K by calculus, re-describe K using the components, and calculate K M×M The eigenvector φ of is given by the following formula:

[0029] K i×j =∫ T a i (t)a j (t)dt=vηv T

[0030] φ=Fv

[0031] By the mean of the eigenvectors and the coefficients of the eigenvectors to calculate the new eigenvectors The new feature vector is converted into a vectorized gait profile cycle through the equation, and the final gait profile sequence is generated through modeling. The formula is as follows:

[0032]

[0033] Among them, ψ represents the eigenvector basis, represents the average gait profile vector, f represents modeling of the gait vector, Represents the newly generated gait profile sequence graph.

[0034] Optionally, in one embodiment of the present invention, in step 14, denoising the reconstructed gait profile image includes:

[0035] The gait profile M1 is obtained using the classic morphological processing algorithm, and the formula is as follows:

[0036] M1=erode(M)

[0037] The gait profile M2 is obtained by performing Mask operation and Contours operation on the obtained gait profile M1. The formula is as follows:

[0038] M2=Contours(Mask(M1))

[0039] Among them, Mask means creating a mask image of the same size as the original image, and Contours means contour extraction;

[0040] The gait profile M3 is obtained by performing a Dilate operation on the gait profile M2, and the gait profile M3 is subjected to a Gaussian process to obtain the final denoised gait profile M4. The formula is as follows:

[0041] M3=dilate(M2)

[0042] M4=Gaussian(M3)

[0043] Among them, dilate means to dilate the gait contour image after removing the ghost, and Gaussian means to use Gaussian filtering to smooth the edges of the gait contour image.

[0044] Optionally, in one embodiment of the present invention, in step 14, inputting the denoised gait profile image into a person detection extraction network to obtain the gait profile features of the pedestrian includes:

[0045] The denoised gait profile is input into the person detection extraction network to obtain multiple feature maps. The person detection extraction network consists of multiple convolution stages. The convolution result of each layer is passed to the next convolution layer. The formula is as follows:

[0046]

[0047] in, represents the convolution operation, H l-1 is the output result of the previous convolutional layer, W i Contains multiple filters, b l Represents the bias shared by the normalized features of different contour maps in the same sequence;

[0048] Calculate the normalized activity of each feature map at a certain spatial position, the formula is as follows:

[0049]

[0050] Among them, a, β, and y all represent adjustable configuration parameters;

[0051] The person detection extraction network is trained through logistic regression, and the predictor is constructed through two layers of neural network layers containing trainable parameters. The formula is as follows:

[0052] L(r)=ReLU(ξ+D a,β -D a,r )

[0053] Among them, a represents the true label, β represents the sample with the same label as the true label a, r represents the sample with a different label from the true label a, ξ represents the boundary distance, the larger the boundary distance, the higher the D a,β The closer the expected distance, the better the a,r The farther the distance is, the outer layer of ReLu represents the activation function.

[0054] Optionally, in one embodiment of the present invention, step 2 further comprises:

[0055] The posture information of the pedestrian is estimated according to the gait profile sequence, and the human joint points are normalized according to the posture coordinate sequence in the posture information. The formula is as follows:

[0056]

[0057] Among them, P i is the coordinate of body joint i, p′ i is the normalized result, P neck is the neck coordinate, H hn It is the height difference between the center of the neck and the center of the hip;

[0058] The character skeleton feature extraction network is used to extract the gait skeleton features of the human skeleton. The formula is as follows:

[0059]

[0060] Among them, f in Represents the feature map with human joint points, f out Represents the output gait skeleton features, A k Represents the adjacency matrix of the joint points in the frame, represents the degree matrix between joint points, with all diagonal elements being zero, w k represents the learnable weight matrix, σ(·) represents the activation function;

[0061] The predictor is constructed by cross entropy loss to classify the gait skeleton features. The formula is as follows:

[0062]

[0063] Among them, y i is the true label of body joint i, p i is the predicted probability of body joint i.

[0064] Optionally, in one embodiment of the present invention, in step 3, performing gait recognition based on the fused pedestrian gait features includes:

[0065] Calculate the feature vector distance between the fused pedestrian gait features and the sample features, sort the calculated multiple feature vector distances, and obtain the similarity result. The calculation formula is:

[0066]

[0067] Among them, D m and D n Represent the minimum and maximum values ​​in the distance, D sort Denotes the distance of the sorted feature vector, D n is the normalized similarity result;

[0068] According to the similarity results, the similarity results of the same sample label are summed and sorted, and multiple results with the highest similarity are selected. The probability distribution function is used to calculate the probability of the sample label to which the fused pedestrian gait features belong, and the gait recognition result is determined based on the probability.

[0069] The gait recognition method based on multimodal feature fusion in low-light scenes of the embodiment of the present invention first obtains a batch of pictures from the collected video to generate an initial detection set, then inputs the initial detection set into the character detection extraction network to extract the gait contour map in the picture and performs noise reduction processing on the gait contour map to improve the clarity of the gait contour map in the low-light scene, then inputs the initial detection set into the character skeleton feature extraction network to generate a skeleton model of the pedestrian in the picture; then inputs the processed gait contour map and skeleton model into the multimodal feature fusion network, and fuses the two logical outputs in the fusion network to obtain the final pedestrian recognition result. The present invention can make full use of various information existing in the picture under low-light conditions, obtain a more comprehensive feature representation, and effectively improve the accuracy, robustness and practicality of gait recognition in low-light scenes.

[0070] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0072] Figure 1 A flowchart of a gait recognition method based on multimodal feature fusion in a low-light scenario provided according to an embodiment of the present invention;

[0073] Figure 2 A flow chart of a gait recognition method according to the present invention;

[0074] Figure 3 This is an architecture diagram of the gait feature extraction network based on multimodal feature fusion of the present invention. DETAILED DESCRIPTION

[0075] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.

[0076] Figure 1 The present invention provides a flowchart of a gait recognition method based on multimodal feature fusion in a low-light scenario according to an embodiment of the present invention.

[0077] like Figure 1 As shown, the gait recognition method based on multimodal feature fusion in low-light scenes includes the following steps:

[0078] Step 1: obtain a gait contour sequence based on a single-person walking video in a low-light scene, remodel the contour sequence based on a partial contour sequence and perform denoising, and input the denoised gait contour into a person detection and extraction network to obtain the pedestrian's gait contour features.

[0079] In an embodiment of the present invention, step 1 further comprises:

[0080] Step 11, processing the CASIA-B dataset to simulate the gait of pedestrians in low-light scenes, and obtaining a single-person walking video in low-light scenes;

[0081] Step 12, using a pedestrian detection and segmentation algorithm to extract and segment the single-person walking image in the single-person walking video into a gait profile sequence;

[0082] Step 13, for the gait profile sequence, by synthesizing gait from partial sequence frames, the characteristics of the gait cycle are modeled as a Gaussian distribution, and the characteristics of the gait cycle are expressed as a continuous function, and a complete cycle gait is generated through modeling;

[0083] Step 14, denoising the reconstructed gait profile image, and inputting the denoised gait profile image into a person detection and extraction network to obtain the gait profile features of the pedestrian.

[0084] In an embodiment of the present invention, step 13 further comprises:

[0085] The dimension of the gait profile data is reduced and features are extracted. The formula is as follows:

[0086] ξλξ T =(SS T )×1 / N

[0087] Among them, ξ represents the common feature vector basis generated by the gait dataset, S represents the shape vector matrix, and N represents the total number of contours in the gait dataset;

[0088] The shape vector u i Project it onto the common eigenvalue basis vector ξ and obtain the corresponding eigenvalue a i , the formula is as follows:

[0089] a i =ξ T u i

[0090] Among them, T represents the transpose of the common feature vector basis generated by the gait dataset, u i represents the shape vector;

[0091] The feature vector a iTransform it into a continuous function, and align the transformed eigenvectors and enhance them into a zero-mean eigenmatrix F, as follows:

[0092] F≡[a1(t),…,a M (t)] m×∞

[0093] Among them, a i (t) represents the feature vector a i The transformed eigenvector, m represents the number of eigenvectors, and ∞ represents that F has infinite dimensions;

[0094] The infinite-dimensional F matrix is ​​converted into a finite-dimensional feature covariance matrix as follows:

[0095] K M×M =F T F

[0096] Among them, K M×M represents a finite-dimensional eigenvector matrix, and F is a characteristic matrix with zero mean;

[0097] Calculate the components of K by calculus, re-describe K using the components, and calculate K M×M The eigenvector φ of is given by the following formula:

[0098] K i×j =∫ T a i (t)a j (t)dt=vηv T

[0099] φ=Fv

[0100] By the mean of the eigenvectors and the coefficients of the eigenvectors to calculate the new eigenvectors The new feature vector is converted into a vectorized gait profile cycle through the equation, and the final gait profile sequence is generated through modeling. The formula is as follows:

[0101]

[0102] Among them, ψ represents the eigenvector basis, represents the average gait profile vector, f represents modeling of the gait vector, Represents the newly generated gait profile sequence graph.

[0103] In an embodiment of the present invention, in step 14, denoising the reconstructed gait profile image includes:

[0104] The gait profile M1 is obtained using the classic morphological processing algorithm, and the formula is as follows:

[0105] M1=erode(M)

[0106] The gait profile M2 is obtained by performing Mask operation and Contours operation on the obtained gait profile M1. The formula is as follows:

[0107] M2=Contours(Mask(M1))

[0108] Among them, Mask means creating a mask image of the same size as the original image, and Contours means contour extraction;

[0109] The gait profile M3 is obtained by performing a Dilate operation on the gait profile M2, and the gait profile M3 is subjected to a Gaussian process to obtain the final denoised gait profile M4. The formula is as follows:

[0110] M3=dilate(M2)

[0111] M4=Gaussian(M3)

[0112] Among them, dilate means to dilate the gait contour image after removing the ghost, and Gaussian means to use Gaussian filtering to smooth the edges of the gait contour image.

[0113] In an embodiment of the present invention, in step 14, the gait profile image after noise reduction is input into the person detection extraction network to obtain the gait profile features of the pedestrian, including:

[0114] The denoised gait profile is input into the person detection extraction network to obtain multiple feature maps. The person detection extraction network consists of multiple convolution stages. The convolution result of each layer is passed to the next convolution layer. The formula is as follows:

[0115]

[0116] in, represents the convolution operation, H l-1 is the output result of the previous convolutional layer, W i Contains multiple filters, b l Represents the bias shared by the normalized features of different contour maps in the same sequence;

[0117] Calculate the normalized activity of each feature map at a certain spatial position, the formula is as follows:

[0118]

[0119] Among them, a, β, and y all represent adjustable configuration parameters;

[0120] The person detection extraction network is trained through logistic regression, and the predictor is constructed through two layers of neural network layers containing trainable parameters. The formula is as follows:

[0121] L(r)=ReLu(ξ+D a,β -D a,r )

[0122] Among them, a represents the true label, β represents the sample with the same label as the true label a, r represents the sample with a different label from the true label a, ξ represents the boundary distance, the larger the boundary distance, the higher the D a,β The closer the expected distance, the better the a,r The farther the distance is, the ReLU in the outer layer represents the activation function.

[0123] Step 2: Input the gait profile sequence into the human skeleton feature extraction network to extract the skeleton features of the pedestrians in the image.

[0124] In an embodiment of the present invention, step 2 further comprises:

[0125] The pedestrian's posture information is estimated based on the gait profile sequence, and the human joint points are normalized according to the posture coordinate sequence in the posture information. The formula is as follows:

[0126]

[0127] Among them, P i is the coordinate of body joint i, p′ i is the normalized result, P neck is the neck coordinate, H hn It is the height difference between the center of the neck and the center of the hip;

[0128] The character skeleton feature extraction network is used to extract the gait skeleton features of the human skeleton. The formula is as follows:

[0129]

[0130] Among them, f in Represents the feature map with human joint points, f out Represents the output gait skeleton features, A k Represents the adjacency matrix of the joint points in the frame, represents the degree matrix between joint points, with all diagonal elements being zero, w k represents the learnable weight matrix, σ(·) represents the activation function;

[0131] The predictor is constructed by cross entropy loss to classify the gait skeleton features. The formula is as follows:

[0132]

[0133] Among them, y i is the true label of body joint i, p i is the predicted probability of body joint i.

[0134] Step 3: Input the pedestrian's gait contour features and skeleton features into the multimodal feature fusion network for feature fusion to obtain the fused pedestrian gait features, and perform gait recognition based on the fused pedestrian gait features.

[0135] In one embodiment of the present invention, in step 3, performing gait recognition based on the fused pedestrian gait features includes:

[0136] Calculate the feature vector distance between the fused pedestrian gait features and the sample features, sort the calculated multiple feature vector distances, and obtain the similarity result. The calculation formula is:

[0137]

[0138] Among them, D m and D n Represent the minimum and maximum values ​​in the distance, D sort Denotes the distance of the sorted feature vector, D n is the normalized similarity result;

[0139] According to the similarity results, the similarity results of the same sample label are summed and sorted, and multiple results with the highest similarity are selected. The probability distribution function is used to calculate the probability of the sample label to which the fused pedestrian gait features belong, and the gait recognition result is determined based on the probability.

[0140] The gait recognition method based on multimodal feature fusion in low-light scenes of the present invention first extracts a gait contour sequence from a single-person walking video, remodels the contour sequence based on the partial contour sequence, then performs noise reduction on the modeled contour, extracts gait features from the noise-reduced gait contour, and simultaneously extracts the original image sequence from the walking video, extracts skeleton features based on the original image sequence, and fuses the two feature extraction networks to form a dual-branch neural network, which combines the features of the skeleton and gait, and further fuses them into the final features of the gait sequence to improve the accuracy and robustness of gait recognition and cope with the challenges brought by practical applications in low-light scenes. Specifically, it includes:

[0141] 1. The present invention uses the RVM algorithm to extract the gait contour sequence of a single-person video as the input representation of the pedestrian gait, re-models the extracted gait contour sequence to obtain a more complete gait contour sequence, and performs noise reduction on the modeled gait contour sequence.

[0142] 2. The present invention uses a 3DCNN network to extract spatiotemporal information and gait features from the denoised gait contour sequence, and improves the accuracy of feature extraction by the neural network through methods such as loss functions.

[0143] 3. The present invention utilizes graph convolutional neural network to extract temporal and spatial features of human skeleton, and constructs an effective predictor through cross entropy loss to classify gait skeleton features.

[0144] 4. The present invention integrates the gait contour feature extraction network and the skeleton feature extraction network into a two-branch neural network, obtains the final robust gait feature by integrating the gait contour feature and the skeleton feature, calculates the similarity between the final gait feature and the sample feature of the gallery set, and outputs the top-k results to realize gait recognition of the person in the video.

[0145] The gait recognition method based on multimodal feature fusion in low-light scenarios of the present invention is described in detail below through a specific embodiment.

[0146] Considering that pedestrian videos shot with ordinary cameras in low-light conditions contain incomplete and unclear information about people, the reconstruction of gait profiles can be used to supplement the information and features of pedestrians in the video. However, there is still some noise in the reconstructed gait profiles, so a noise reduction algorithm is designed to reduce the noise of the virtual shadows in the gait profiles to obtain a clearer gait profile. At the same time, considering the cross-viewing angle, the side gait profile is still unavailable after reconstruction and noise reduction. The pedestrian skeleton features are used to more accurately identify the pedestrians at the measured angle. Therefore, based on the above situation, a dual-branch neural network is designed to obtain the final gait features by fusing the gait profile features and skeleton features, so as to achieve more accurate identification of pedestrians in low-light scenes using ordinary cameras.

[0147] like Figure 2 and Figure 3 As shown, the gait recognition method based on multimodal feature fusion in low-light scenes includes the following steps:

[0148] Step 1: Process the CASIA-B dataset to simulate pedestrian gait in low-light scenes.

[0149] Step 2: Use the pedestrian detection and segmentation algorithm Robust Video Matting algorithm to extract and segment the single person walking image in the gait video processed in step 1 into a gait contour sequence.

[0150] Step 3: For the gait contour sequence obtained in step 2, the corresponding features of the gait cycle are modeled as Gaussian distribution by synthesizing gait from a small number of sequence frames, and these features are expressed as continuous functions. Finally, a relatively complete periodic gait is generated through modeling.

[0151] Step 4: Extract the spatiotemporal information and features of the gait profile data obtained in step 3 through the 3DCNN network, and improve the accuracy of feature extraction by the neural network through methods such as loss function.

[0152] Step 5: Create a gait dataset in a real low-light scene, label the dataset, and repeat steps 2 and 3 to obtain the gait profile of the self-made dataset.

[0153] Step 6: Design an algorithm to perform noise reduction on the gait profile image obtained in step 5 to improve the accuracy of gait recognition.

[0154] Step 7: Extract skeleton features from the self-made data set in step 6, and use networks such as GaitGraph to extract and train the skeleton contour of the character.

[0155] Step 8: Input the denoised gait contour image obtained in step 6 and the skeleton contour in step 7 into the pre-trained multimodal feature fusion network for feature fusion to realize gait recognition of the characters in the self-made dataset in low-light scenes.

[0156] Step 9: Calculate the feature similarity between the final gait feature obtained in step 8 and multiple sample features in the gallery set, and output the sample ID with the highest similarity to the target person.

[0157] Specifically, the gait profile sequence extraction in step 2 includes the following steps:

[0158] Step 2-1: Crop the pedestrian walking video obtained in step 1 and step 5, and crop the video into image sequences of different lengths;

[0159] Step 2-2: Normalize the image obtained in step 2-1, and crop the image obtained in step 2-1 to obtain an image with a length and width of 64 and 60;

[0160] Step 2-3: extracting the character outline of each frame image obtained in step 2-2 through the RVM network;

[0161] Specifically, the modeling method of the gait profile sequence in step 3 includes the following steps:

[0162] Step 3-1: First, reduce the dimension of the gait profile data through PCA, and then extract specific features. The formula is as follows:

[0163] ξλξT =(SS T )×1 / N

[0164] Among them, ξ represents the common feature vector basis generated by the entire gait dataset, S represents the shape vector matrix obtained by PCA, and N represents the total number of contours in the gait dataset;

[0165] Step 3-2: Transform the shape vector u i Project it onto the common feature basis vector ξ and obtain the corresponding feature mode a i , the formula is as follows:

[0166] a i =ξ T u i

[0167] Among them, ξ T represents the transpose of the common feature vector basis generated by the entire gait dataset, u i represents the shape vector;

[0168] Step 3-3: Transform the feature vector a i Transform it into a basic continuous function, and align the transformed eigenvectors and enhance them into a zero-mean eigenmatrix F, as follows:

[0169] F≡[a1(t),…,a M (t)] m×∞

[0170] Among them, a i (t) represents the feature vector a i The transformed eigenvector, m represents the number of eigenvectors, and ∞ represents that F has infinite dimensions.

[0171] Step 3-4: Convert the infinite-dimensional F matrix obtained in step 3-3 into a finite-dimensional feature covariance matrix. The formula is as follows:

[0172] K M×M =F T F

[0173] Among them, K M×M represents a finite-dimensional eigenvector matrix, and F is a characteristic matrix with zero mean;

[0174] Step 3-5: Calculate the components of K by calculus, re-describe K using the components, and calculate K M×M The eigenvector φ of is given by the following formula:

[0175] K i×j =∫ T a i (t)a j (t)dt=vηvT

[0176] φ=Fv

[0177] Step 3-6: By taking the mean of the eigenvectors and the coefficients of the eigenvectors to calculate the new eigenvectors These new feature vectors are converted into vectorized gait profile cycles through equations, and the final gait profile sequence is generated through modeling. The formula is as follows:

[0178]

[0179] Among them, ψ represents the eigenvector basis of PCA, represents the average gait profile vector, f represents the modeling of the gait vector, Represents the newly generated gait profile sequence graph.

[0180] Specifically, the method for extracting the features of the gait profile graph in step 4 comprises the following steps:

[0181] Step 4-1: Obtain the gait profile image after noise reduction through step 7;

[0182] Step 4-2: Input the gait profile image into the CNN network, where the CNN network is composed of multiple convolution stages, and the convolution result of each layer is passed to the next convolution layer. The formula is as follows:

[0183]

[0184] in, represents the convolution operation, H l-1 is the output result of the convolutional layer of the previous layer, H0 is the original input data, W i Contains multiple filters, b l Represents the bias shared by the normalized features of different profiles in the same sequence.

[0185] Step 4-3: Calculate the normalized activity of each feature map at a certain spatial position, the formula is as follows:

[0186]

[0187] Among them, a, β, and y all represent adjustable configuration parameters.

[0188] Step 4-4: Train the entire network through logistic regression, and form an effective predictor through two layers of neural network layers containing trainable parameters. The formula is as follows:

[0189] L(r)=ReLU(ξ+D a,β -D a,r )

[0190] Among them, a represents the true label, β represents the sample with the same label as the true label a, r represents the sample with a different label from the true label a, and ξ represents the boundary distance. If the value is larger, it means that D a,β The closer the expected distance, the better the a,r The outer ReLu represents the activation function.

[0191] Specifically, the denoising method of the gait profile image in step 6 specifically comprises the following steps:

[0192] Step 6-1: Use the classic morphological processing algorithm Erode to obtain the contour M1. The formula is as follows:

[0193] M1=erode(M)

[0194] Step 6-2: Obtain M2 by performing Mask and Contours operations on M1 obtained in step 6-1. The formula is as follows:

[0195] M2=Contours(Mask(M1))

[0196] Among them, Mask means creating a mask image of the same size as the original image, and Contours means contour extraction.

[0197] Step 6-3: Perform a Dilate operation on the contour map M2 obtained in step 6-2 to obtain the contour map M3, and perform Gaussian processing on M3 to obtain the final denoised contour map M4. The formula is as follows:

[0198] M3=dilate(M2)

[0199] M4=Gaussian(M3)

[0200] Among them, dilate means to dilate the contour image after removing the ghost, and Gaussian means to use Gaussian filtering to smooth the edges of the contour image.

[0201] Specifically, step 7, the skeleton feature extraction method of the self-made data set includes the following steps:

[0202] Step 7-1: Estimate posture information based on the gait profile sequence obtained in step 2;

[0203] Step 7-2: Extract the posture coordinate sequence in the posture information obtained in step 7-1 and perform preprocessing to normalize the human body joint points. The formula is as follows:

[0204]

[0205] Among them, P iis the coordinate of body joint i, p′ i is the normalized result, P neck is the neck coordinate, H n It is the height difference between the center of the neck and the center of the hip;

[0206] Step 7-3: Use graph convolutional neural network to extract the temporal and spatial features of the human skeleton. The formula is as follows:

[0207]

[0208] Among them, f in Represents the feature map with human joint points, f out Represents the final skeleton feature output, A k Represents the adjacency matrix of the joint points in the frame, represents the degree matrix between joint points, whose diagonal elements are all zero, w k represents the learnable weight matrix, σ(·) represents the activation function;

[0209] Step 7-4: Use cross entropy loss to construct an effective predictor to classify the gait skeleton features obtained in step 7-3. The formula is as follows:

[0210]

[0211] Among them, y i is the true label of sample i, p i is the predicted probability of sample i.

[0212] Specifically, step 8 multimodal feature fusion includes the following steps: designing a two-branch neural network, merging the gait contour feature extraction network obtained in step 4 and the skeleton contour feature extraction network obtained in step 7-2 as different branches into a two-branch network, and finally aggregating the two types of features into the final robust gait feature.

[0213] Specifically, step 9 calculates the feature similarity between the gait feature and the sample feature in the sample space, and specifically includes the following steps:

[0214] Step 9-1: Calculate the distance between feature vectors based on the final gait features obtained in step 8 and the sample features in the gallery set;

[0215] Step 9-2: Sort the feature vector distances calculated in step 9-1 to obtain the similarity result. The formula is as follows:

[0216]

[0217] Among them, D m and D nRepresent the minimum and maximum values ​​in the distance, D sort Denotes the distance of the sorted feature vector, D n Then it is the normalized similarity result;

[0218] Step 9-3: Based on the similarity obtained in step 9-2, sum and sort the similarity results of the same sample label, and take the top-5 results with the highest similarity;

[0219] Step 9-4: Based on the top-5 results obtained in step 9-3, use the probability distribution function to calculate the probability of the sample label to which the contour sequence belongs, and output the gait recognition result.

[0220] The gait recognition method based on multimodal feature fusion in low-light scenes proposed in an embodiment of the present invention can realize relatively accurate recognition of pedestrians in low-light scenes using ordinary cameras, mainly through reconstruction of the pedestrian gait contour map and noise reduction processing of the modeled gait contour map, so as to obtain a clearer gait contour map, fully explore and utilize the information of the pedestrian gait contour map, use skeleton feature extraction to make up for the defect that the angle-measured gait contour is still unusable after reconstruction and noise reduction, flexibly fuse the skeleton feature and gait contour feature to form the final robust gait feature, and effectively improve the robustness and accuracy of gait recognition in low-light scenes.

[0221] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.

[0222] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.

[0223] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present invention belong.

Claims

1. A gait recognition method based on multimodal feature fusion in low-light scenes, characterized in that: The following steps are involved: Step 1: obtain a gait contour sequence based on a single person walking video in a low-light scene, remodel the contour sequence based on a partial contour sequence and perform noise reduction, and input the denoised gait contour into a person detection and extraction network to obtain the pedestrian's gait contour features; Step 2, inputting the gait profile sequence into a human skeleton feature extraction network to extract skeleton features of pedestrians in the image; Step 3: Input the pedestrian's gait profile features and the skeleton features into a multimodal feature fusion network for feature fusion to obtain fused pedestrian gait features, and perform gait recognition based on the fused pedestrian gait features.

2. The method according to claim 1, characterized in that The step 1 further comprises: Step 11, processing the CASIA-B dataset to simulate the gait of pedestrians in low-light scenes, and obtaining a single-person walking video in low-light scenes; Step 12, using a pedestrian detection and segmentation algorithm to extract and segment the single-person walking image in the single-person walking video into a gait profile sequence; Step 13, for the gait profile sequence, by synthesizing gait from partial sequence frames, the characteristics of the gait cycle are modeled as a Gaussian distribution, and the characteristics of the gait cycle are expressed as a continuous function, and a complete cycle gait is generated through modeling; Step 14, denoising the reconstructed gait profile image, and inputting the denoised gait profile image into a person detection and extraction network to obtain the gait profile features of the pedestrian.

3. The method according to claim 1, characterized in that Step 13 further includes: The dimension of the gait profile data is reduced and features are extracted. The formula is as follows: zlz T =(SS T )×1 / N Among them, ζ represents the common feature vector basis generated by the gait dataset, S represents the shape vector matrix, and N represents the total number of contours in the gait dataset; The shape vector u i Project it onto the common eigenvalue basis vector ξ and obtain the corresponding eigenvalue a i , the formula is as follows: a i =ξ T you i Among them, ξ T represents the transpose of the common feature vector basis generated by the gait dataset, u i represents the shape vector; The feature vector a i Transform it into a continuous function, and align the transformed eigenvectors and enhance them into a zero-mean eigenmatrix F, as follows: F≡[a1(t),…,a M (t)] m×∞ Among them, a i (t) represents the feature vector a i The transformed eigenvector, m represents the number of eigenvectors, and ∞ represents that F has infinite dimensions; The infinite-dimensional F matrix is ​​converted into a finite-dimensional feature covariance matrix as follows: K M×M =F T F Among them, K M×M represents a finite-dimensional eigenvector matrix, and F is a characteristic matrix with zero mean; Calculate the components of K by calculus, re-describe K using the components, and calculate K M×M The eigenvector φ of is given by the following formula: K i×j =∫ T a i (t)a j (t)dt=vηv T φ=Fv By the mean of the eigenvectors and the coefficients of the eigenvectors to calculate the new eigenvectors The new feature vector is converted into a vectorized gait profile cycle through the equation, and the final gait profile sequence is generated through modeling. The formula is as follows: Among them, ψ represents the eigenvector basis, represents the average gait profile vector, f represents modeling of the gait vector, Represents the newly generated gait profile sequence graph.

4. The method according to claim 1, characterized in that: In step 14, denoising the reconstructed gait profile image includes: The gait profile M1 is obtained using the classic morphological processing algorithm, and the formula is as follows: M1=erode(M) The gait profile M2 is obtained by performing Mask operation and Contours operation on the obtained gait profile M1. The formula is as follows: M2=Contours(Mask(M1)) Among them, Mask means creating a mask image of the same size as the original image, and Contours means contour extraction; The gait profile M3 is obtained by performing a Dilate operation on the gait profile M2, and the gait profile M3 is subjected to a Gaussian process to obtain the final denoised gait profile M4. The formula is as follows: M3=dilate(M2) M4=Gaussian(M3) Among them, dilate means to dilate the gait contour image after removing the ghost, and Gaussian means to use Gaussian filtering to smooth the edges of the gait contour image.

5. The method according to claim 1, characterized in that In step 14, the denoised gait profile image is input into the person detection extraction network to obtain the pedestrian's gait profile features including: The denoised gait profile is input into the person detection extraction network to obtain multiple feature maps. The person detection extraction network consists of multiple convolution stages. The convolution result of each layer is passed to the next convolution layer. The formula is as follows: in, represents the convolution operation, H l-1 is the output result of the previous convolutional layer, W i Contains multiple filters, b l Represents the bias shared by the normalized features of different contour maps in the same sequence; Calculate the normalized activity of each feature map at a certain spatial position, the formula is as follows: Among them, a, β, and y all represent adjustable configuration parameters; The person detection extraction network is trained through logistic regression, and the predictor is constructed through two layers of neural network layers containing trainable parameters. The formula is as follows: L(r)=ReLU(ξ+D a,β -D a,r ) Among them, a represents the true label, β represents the sample with the same label as the true label a, r represents the sample with a different label from the true label a, ξ represents the boundary distance, the larger the boundary distance, the higher the D a,β The closer the expected distance, the better the a,r The farther the distance is, the outer layer of ReLu represents the activation function.

6. The method according to claim 1, characterized in that Step 2 further includes: The posture information of the pedestrian is estimated according to the gait profile sequence, and the human joint points are normalized according to the posture coordinate sequence in the posture information. The formula is as follows: Among them, P i is the coordinate of body joint i, p′ i is the normalized result, P neck is the neck coordinate, H hn It is the height difference between the center of the neck and the center of the hip; The character skeleton feature extraction network is used to extract the gait skeleton features of the human skeleton. The formula is as follows: Among them, f in Represents the feature map with human joint points, f out Represents the output gait skeleton features, A k Represents the adjacency matrix of the joint points in the frame, represents the degree matrix between joint points, with all diagonal elements being zero, w k represents the learnable weight matrix, σ(·) represents the activation function; The predictor is constructed by cross entropy loss to classify the gait skeleton features. The formula is as follows: Among them, y i is the true label of body joint i, p i is the predicted probability of body joint i.

7. The method according to claim 1, characterized in that In step 3, performing gait recognition based on the fused pedestrian gait features includes: Calculate the feature vector distance between the fused pedestrian gait features and the sample features, sort the calculated multiple feature vector distances, and obtain the similarity result. The calculation formula is: Among them, D m and D n Represent the minimum and maximum values ​​in the distance, D sort Denotes the distance of the sorted feature vector, D n is the normalized similarity result; According to the similarity results, the similarity results of the same sample label are summed and sorted, and multiple results with the highest similarity are selected. The probability distribution function is used to calculate the probability of the sample label to which the fused pedestrian gait features belong, and the gait recognition result is determined based on the probability.

Citation Information

Patent Citations

  • Gait recognition method based on modal fusion

    CN111428658A

  • Homomorphic heterogeneous data feature extraction method for gait recognition

    CN116403276A

  • Heterogeneous data collaborative gait recognition method based on Gaussian distribution modeling

    CN118865450A

  • Gait recognition method

    KR100824757B1

Cited By

  • Human body gait recognition method and system based on multivariate features

    CN121121851A

  • Gait recognition method, device and equipment and storage medium

    CN121214552A