A gait recognition method, device and equipment based on multi-modal and a storage medium

By using multimodal image acquisition and processing technology, gait features of pedestrians in different modalities are extracted, which solves the problem of low accuracy in gait recognition in existing technologies and achieves higher accuracy in gait feature recognition.

CN116721438BActive Publication Date: 2026-03-31WATRIX TECH CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies use only RGB video images as the basic image data in pedestrian gait recognition, resulting in low accuracy of gait feature recognition results.

Method used

A multimodal input method is adopted, using various image acquisition devices (such as RGB cameras, depth cameras, infrared cameras, event cameras, and LiDAR) to acquire walking video images of pedestrians. Silhouette image sequences of the same pedestrian are extracted through image target detection and instance segmentation, and then input into a pre-trained gait recognition model for feature recognition.

Benefits of technology

The accuracy of gait feature recognition results has been improved by combining feature recognition results from multiple modal data, thereby enhancing the accuracy of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116721438B_ABST
    Figure CN116721438B_ABST
Patent Text Reader

Abstract

The application provides a gait recognition method and device based on multiple modes, equipment and a storage medium. The method comprises: in the same space range, using multiple image acquisition devices of different modes, simultaneously collecting walking video images of different pedestrians in the space range to obtain walking video images of each mode; performing image target detection on pedestrians included in the walking video images of each mode, and separating silhouette image sequences of the same pedestrian from the walking video images of each mode according to the image target detection result; inputting the multiple silhouette image sequences of different modes corresponding to the same pedestrian into a pre-trained gait recognition model to output a gait feature recognition result for the pedestrian. In this way, the application is based on the multiple mode input mode, so that the gait feature recognition result output contains different gait features of the same pedestrian in different mode data, effectively improving the accuracy of the gait feature recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of gait detection technology, and more specifically, to a multimodal gait recognition method, apparatus, device, and storage medium. Background Technology

[0002] Currently, existing technologies for pedestrian gait recognition typically only collect RGB (red, green, and blue) video images of different pedestrians walking as the basic image data for gait feature recognition. Therefore, according to existing gait recognition methods, only the gait features exhibited by different pedestrians in this single modality of RGB video image data can be identified, resulting in low accuracy of the obtained gait feature recognition results. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a gait recognition method, apparatus, device and storage medium based on multimodal input. Based on multimodal input, the silhouette image sequence of the same pedestrian in multiple different modalities is used as the model input data of the gait recognition model at the same time. This makes the output gait feature recognition result contain the different gait features of the same pedestrian in different modal data, which effectively improves the accuracy of gait feature recognition results.

[0004] In a first aspect, embodiments of this application provide a gait recognition method based on multimodal modes, the gait recognition method comprising:

[0005] Within the same spatial range, multiple image acquisition devices of different modalities are used to simultaneously acquire walking video images of different pedestrians within that spatial range, thus obtaining walking video images of each modality;

[0006] Image target detection is performed on pedestrians included in the walking video images of each modality, and silhouette image sequences belonging to the same pedestrian are separated from the walking video images of each modality based on the image target detection results;

[0007] A sequence of silhouette images of multiple different modalities corresponding to the same pedestrian is input into a pre-trained gait recognition model, and the output is the gait feature recognition result for the pedestrian; wherein, the gait feature recognition result represents the splicing result of the gait features of multiple different modalities corresponding to the pedestrian.

[0008] Secondly, embodiments of this application provide a gait recognition device based on multimodal modes, the gait recognition device comprising:

[0009] The data acquisition module is used to simultaneously acquire walking video images of different pedestrians within the same spatial range using multiple image acquisition devices of different modalities, thereby obtaining walking video images of each modality.

[0010] The image processing module is used to perform image target detection on pedestrians included in the walking video images of each modality, and to separate the silhouette image sequence belonging to the same pedestrian from the walking video images of each modality based on the image target detection results;

[0011] The feature recognition module is used to input a sequence of silhouette images of multiple different modalities corresponding to the same pedestrian into a pre-trained gait recognition model and output a gait feature recognition result for the pedestrian; wherein, the gait feature recognition result represents the splicing result of the gait features of multiple different modalities corresponding to the pedestrian.

[0012] Thirdly, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the gait recognition method described above.

[0013] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the gait recognition method described above.

[0014] The technical solutions provided by the embodiments of this application may include the following beneficial effects:

[0015] This application provides a multimodal gait recognition method, apparatus, device, and storage medium. Within the same spatial range, multiple image acquisition devices of different modalities are used to simultaneously acquire walking video images of different pedestrians within that spatial range, obtaining walking video images of each modality. Image target detection is performed on the pedestrians included in each modality of the walking video images, and based on the image target detection results, silhouette image sequences belonging to the same pedestrian are separated from each modality of the walking video images. The multiple silhouette image sequences of the same pedestrian corresponding to the same modality are input into a pre-trained gait recognition model, outputting a gait feature recognition result for that pedestrian. Thus, this application, based on a multimodal input approach, uses multiple silhouette image sequences of the same pedestrian in different modalities simultaneously as model input data for the gait recognition model. This ensures that the output gait feature recognition result contains the different gait features exhibited by the same pedestrian in different modal data, effectively improving the accuracy of the gait feature recognition result.

[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a multimodal gait recognition method provided in an embodiment of this application is shown.

[0019] Figure 2 This illustration shows a schematic diagram of the model structure of a gait recognition model provided in an embodiment of this application;

[0020] Figure 3 The diagram shows a flowchart of a training method for a gait recognition model provided in an embodiment of this application.

[0021] Figure 4 A flowchart illustrating a gait feature retrieval method provided in an embodiment of this application is shown;

[0022] Figure 5 A schematic diagram of the structure of a multimodal gait recognition device provided in an embodiment of this application is shown;

[0023] Figure 6 This is a schematic diagram of the structure of a computer device 600 provided in an embodiment of this application. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0025] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0026] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0027] Currently, existing technologies for pedestrian gait recognition typically only collect RGB video images of different pedestrians walking as the basic image data for gait feature recognition. Therefore, according to existing gait recognition methods, only the gait features exhibited by different pedestrians in this single modality of image data, such as RGB video images, can be identified, resulting in low accuracy of the gait feature recognition results.

[0028] Based on this, embodiments of this application provide a method, apparatus, device, and storage medium for gait recognition. Based on a multimodal input approach, silhouette image sequences of multiple different modalities of the same pedestrian are simultaneously used as model input data for a gait recognition model. This results in the output gait feature recognition results containing different gait features exhibited by the same pedestrian in different modal data, effectively improving the accuracy of gait feature recognition results.

[0029] The following is a detailed description of a multimodal gait recognition method, apparatus, device, and storage medium provided in the embodiments of this application.

[0030] Reference Figure 1 As shown, Figure 1 The diagram illustrates a flowchart of a multimodal gait recognition method provided in an embodiment of this application, wherein the gait recognition method includes steps S101-S103; specifically:

[0031] S101, within the same spatial range, using multiple image acquisition devices of different modalities, simultaneously acquire walking video images of different pedestrians within that spatial range, and obtain walking video images of each modality.

[0032] Here, each modal image acquisition device is used to selectively acquire walking video images of the same target object (i.e., different pedestrians within the same spatiotemporal range) in a specific modal type; wherein, the acquired walking video images of each modal can include multiple consecutive image frames, that is, each modal walking video image is equivalent to an image sequence of a specific modal type.

[0033] It should be noted that in step S101 above, it is only necessary to ensure that the multiple image acquisition devices of different modalities acquire pedestrian walking video images within the same spatiotemporal range (i.e., simultaneously acquire images of the same group of pedestrians within the same spatial range) to obtain multiple walking video images of the same target object (which can be one pedestrian or multiple pedestrians). This application embodiment does not impose any limitations on the specific shooting angle, specific device type, specific shooting duration (which is also equivalent to the specific number of images included in the walking video image), and specific number of pedestrians captured by different image acquisition devices.

[0034] Specifically, in the embodiments of this application, in order to obtain walking video images of various different modalities, the image acquisition device used may include the following different types of example devices:

[0035] 1. RGB camera; Among them, an RGB camera can capture walking video images with RGB modal type.

[0036] 2. Depth camera; Among them, a depth camera can be used to capture walking video images with depth information modality (i.e., walking video images with depth information).

[0037] 3. Infrared camera; among which, an infrared camera can capture walking video images in infrared mode.

[0038] 4. Event camera; The event camera is a bionic sensor with a microsecond response time. It generates an event by detecting the brightness change of each pixel. In other words, the event camera can capture walking video images with a modal type of pixel brightness change (i.e., the walking video images contain pixel brightness change information).

[0039] 5. LiDAR; among which, LiDAR can capture walking video images with a mode of light reflection (i.e., walking video images contain light reflection information).

[0040] It should be noted that in step S101 above, the image acquisition devices of different modalities may include at least two of the devices listed above (e.g., an RGB camera and a depth camera, or an RGB camera, a depth camera, and an infrared camera, etc.). That is, in step S101 above, at least two different modalities of walking video images can be acquired (e.g., walking video images of RGB video image type and walking video images with depth information, or walking video images of RGB video image type, walking video images with depth information, and walking video images of infrared type, etc.). This application embodiment does not limit the specific number of devices, the specific combination method, or the specific modal type corresponding to the acquired walking video images used in step S101 above.

[0041] S102, perform image target detection on pedestrians included in the walking video images of each modality, and based on the image target detection results, separate the silhouette image sequence belonging to the same pedestrian from the walking video images of each modality.

[0042] Here, based on the content of step S101 above, it is known that each mode of walking video image is equivalent to a specific mode type of image sequence; at this time, for each mode type of image sequence (i.e. each mode of walking video image), in step S102, with pedestrians as the detection target, image target detection is performed on each frame of the image sequence, and the target image region belonging to pedestrians and other image regions not belonging to pedestrians in each frame of the image can be determined (i.e., the above image target detection result is obtained).

[0043] It should be noted that in step S102 above, the YOLOv5 object detection algorithm (or YOLOv5 object detection model) can be used to perform image object detection on pedestrians included in each modality of walking video images; in addition to the YOLOv5 object detection algorithm, the YOLOv4 object detection algorithm, the SSD (Single Shot MultiBox Detector) object detection algorithm, etc. can also be used to implement the above image object detection steps. This application embodiment does not limit the specific image object detection method used in step S102 above.

[0044] Here, for each modality of image sequence (i.e., walking video images of each modality), after obtaining the image target detection results of the image sequence, the target image region belonging to the pedestrian in each frame of the image has been marked in the image target detection results. Therefore, the silhouette image sequence belonging to the same pedestrian (i.e., the silhouette image sequence composed of the specific outline of the same pedestrian) can be separated from the walking video images of each modality by performing instance segmentation on the image target detection results.

[0045] Specifically, for each mode of walking video image, after obtaining the image target detection results for that mode of walking video image, instance segmentation can be performed according to the following steps a1-a2:

[0046] Step a1: For each mode of walking video image, perform instance segmentation on the image target detection results of the walking video image of that mode to obtain silhouette images of different pedestrians corresponding to that mode and pedestrian labels corresponding to each frame of the silhouette image.

[0047] Here, as an optional embodiment, the HTC (Hybrid Task Cascade) segmentation model can be used to perform instance segmentation on the image target detection results of walking video images of each modality. Based on the image target detection results, the specific contour of each pedestrian can be extracted from each frame of the walking video image, and the extracted specific contours of different pedestrians can be used as silhouette images of different pedestrians corresponding to that modality.

[0048] Specifically, when performing instance segmentation, the HTC segmentation model also labels the silhouette images of different pedestrians (i.e., the pedestrian labels mentioned above), where the same pedestrian corresponds to the same label; based on this, according to the instance segmentation results, for each silhouette image extracted above, the pedestrian label corresponding to each silhouette image can also be obtained at the same time.

[0049] It should be noted that, in addition to the HTC segmentation model, other different types of segmentation models can also be used to complete the above instance segmentation steps. This application does not limit the specific type of segmentation model used for instance segmentation.

[0050] Step a2: Based on the pedestrian tag corresponding to each silhouette image, select multiple silhouette images corresponding to the same pedestrian tag from the silhouette images of different pedestrians corresponding to this modality.

[0051] In practical applications, since the walking video images are captured simultaneously by different pedestrians within the same spatial range, the number of specific pedestrians included in each frame of the walking video image is uncertain (it may be one or more). Based on this, after completing the instance segmentation step, it is also necessary to determine the specific silhouette image belonging to the same pedestrian based on the pedestrian tag corresponding to each silhouette image.

[0052] For example, taking an RGB type walking video image as an example, the RGB type walking video image includes 20 frames of images. By performing instance segmentation on the image target detection results of these 20 frames, a total of 32 RGB type silhouette images are obtained. At this time, taking pedestrian label 'a' as an example, according to the pedestrian label corresponding to each of these 32 RGB type silhouette images, 4 silhouette images corresponding to pedestrian label 'a' can be selected. It is determined that these 4 selected silhouette images belong to the same pedestrian (i.e., the pedestrian represented by pedestrian label 'a') and are RGB type silhouette images.

[0053] S103, input the sequence of silhouette images of multiple different modalities corresponding to the same pedestrian into the pre-trained gait recognition model, and output the gait feature recognition result for the pedestrian.

[0054] Here, the gait feature recognition result described above represents the stitching result of gait features of multiple different modalities corresponding to the pedestrian. That is, in this embodiment, the silhouette image sequence of multiple different modalities corresponding to the same pedestrian is used as the model input data of the gait recognition model (i.e., the gait recognition model supports multimodal input), so that the output gait feature recognition result contains the different gait features of the same pedestrian in different modal data. Therefore, compared with the existing gait recognition method that extracts gait recognition features from single modal data, this embodiment can effectively improve the accuracy of gait feature recognition results.

[0055] Specifically, the gait recognition model mentioned above includes multiple channels, which do not interfere with each other. The gait recognition model can perform the same feature extraction processing on the input data in each channel. Based on this, in this embodiment, the specific input method of the above model input data (i.e., the silhouette image sequence of multiple different modalities corresponding to the same pedestrian) has the following two different optional implementation methods:

[0056] Option 1: Mix together multiple silhouette image sequences of different modalities corresponding to the same pedestrian as complete gait recognition data for that pedestrian, and input them into the gait recognition model through a single-channel input method.

[0057] Here, when inputting the model according to the optional implementation method 1 above, step S103 can be implemented in the manner of step b1:

[0058] Step b1: Simultaneously input the silhouette image sequences of the same pedestrian corresponding to the multiple different modalities into the same channel of the gait recognition model, and use the gait recognition model to perform mixed feature extraction on the silhouette image sequences of the multiple different modalities in the channel to output the gait feature recognition result for the pedestrian.

[0059] It should be noted that this application does not impose any limitations on the specific number of channels included in the gait recognition model or the specific channels used in the actual input of the above model data.

[0060] Option 2: Input the silhouette image sequences of multiple different modalities corresponding to the same pedestrian into different channels of the gait recognition model, so as to input them into the gait recognition model through a multi-channel input method.

[0061] Here, when inputting the model according to the optional implementation method 2 described above, step S103 can be implemented in the following steps c1-c2:

[0062] Step c1: Input the silhouette image sequences of the same pedestrian corresponding to the various different modalities into different channels of the gait recognition model. The gait recognition model performs independent feature extraction on the silhouette image sequence in each channel to obtain the gait features of each modality corresponding to the pedestrian.

[0063] Here, unlike the optional implementation method 1 above, each channel corresponds to the input of a silhouette image sequence of a pedestrian's modality, and each channel corresponds to the output of a gait feature of a pedestrian's modality.

[0064] It should be noted that in optional implementation 2, the feature extraction processing of the silhouette image sequence in each channel by the gait recognition model is the same as the feature extraction processing in a single channel in optional implementation 1. However, due to the different specific image data input in the channel, the channel in optional implementation 1 can directly output the gait feature recognition result of the pedestrian, while the channel in optional implementation 2 will only output the gait feature of one mode of the pedestrian.

[0065] Step c2: Perform gait feature concatenation on the gait features of each modality corresponding to the pedestrian, and output the concatenation result as the gait feature recognition result for the pedestrian.

[0066] Here, in optional implementation 2, by splicing the gait features of each modality output by each channel, the same gait feature recognition result as in optional implementation 1 can be obtained (that is, the output gait feature recognition result also contains the different gait features of the same pedestrian in different modal data).

[0067] The specific model structure and training method of the above gait recognition model will be explained in detail below:

[0068] Regarding the specific model structure of the aforementioned gait recognition model, Figure 2 This illustration shows a schematic diagram of the model structure of a gait recognition model provided in an embodiment of this application, as shown below. Figure 2 As shown, the gait recognition model includes at least: a first feature extractor 201, a second feature extractor 202, and an ensemble pooling layer; wherein, the first feature extractor 201 and the second feature extractor 202 have the same structure, and each feature extractor may include multiple convolutional modules (e.g., ...). Figure 2 The first feature extractor 201 and the second feature extractor 202 shown can both be composed of 3 convolutional modules. Each convolutional module can contain 2 CNN (Convolutional Neural Network) convolutional layers and 1 pooling layer (not shown in the figure).

[0069] It should be noted that, Figure 2 The specific model structure shown is only used to illustrate the gait recognition model in the embodiments of this application. The embodiments of this application do not limit the specific number of feature extractors included in the gait recognition model, or the specific number of convolutional blocks included in each feature extractor.

[0070] Specifically, with Figure 2 Taking the model structure shown as an example, a sequence of silhouette images of the same pedestrian in multiple different modalities is used as... Figure 2 The model input data shown above, and step S103 can be specifically executed according to the following steps d1-d7:

[0071] Step d1: For the silhouette image sequences of the same pedestrian corresponding to the multiple different modalities, the first feature extractor in the gait recognition model performs frame-level feature extraction on the silhouette image sequence of each modality to obtain the first frame-level feature sequence of each modality.

[0072] Specifically, taking the infrared silhouette image sequence corresponding to pedestrian A, which includes 5 frames of silhouette images of pedestrian A, as an example, these 5 frames of silhouette images of pedestrian A are input into the first feature extractor 201. The first feature extractor 201 performs frame-level feature extraction on each frame of silhouette image of pedestrian A, and can extract the frame-level features corresponding to each frame of silhouette image of pedestrian A. At this time, the first feature extractor 201 outputs the frame-level features corresponding to each frame of silhouette image of pedestrian A, and thus obtains the first frame-level feature sequence of infrared type of pedestrian A.

[0073] Here, the gait recognition model performs the same feature extraction process for each modality of silhouette image sequence. That is, the first feature extractor can perform frame-level feature extraction for each modality of silhouette image sequence corresponding to the same pedestrian in the manner described in the example above. The repetitions will not be repeated here.

[0074] Step d2: Through the ensemble pooling layer in the gait recognition model, the first frame-level feature sequence of each modality is aggregated to obtain the first ensemble-level feature of each modality.

[0075] Specifically, as an optional embodiment, the ensemble pooling layer can aggregate the first frame-level feature sequences of each modality according to the statistical function shown below:

[0076] G(x)=ax(x)+ean(x)+edian(x);

[0077] Where x represents the first frame-level feature sequence of any modality corresponding to the same pedestrian;

[0078] max(·) represents the max pooling function;

[0079] mean(·) represents the average pooling function;

[0080] median(·) represents the median function;

[0081] G(·) represents the first set-level feature obtained after aggregating the first frame-level feature sequence x.

[0082] For example, taking the silhouette image sequence of pedestrian A in three different modalities (modality a, modality b, and modality c) as the model input data, the first feature extractor can output the first frame-level feature sequences of the three different modalities respectively: x1, x2, and x3. Among them, the first frame-level feature sequence x1 of modality a includes frame-level features corresponding to 5 silhouette images, the first frame-level feature sequence x2 of modality b includes frame-level features corresponding to 4 silhouette images, and the first frame-level feature sequence x3 of modality c includes frame-level features corresponding to 6 silhouette images. At this time, the ensemble pooling layer can aggregate the frame-level features corresponding to the above 5 silhouette images into the first ensemble feature of modality a, aggregate the frame-level features corresponding to the above 4 silhouette images into the first ensemble feature of modality b, and aggregate the frame-level features corresponding to the above 6 silhouette images into the first ensemble feature of modality c.

[0083] Step d3: For the first convolution processing result of the silhouette image sequence of each modality, the second feature extractor in the gait recognition model performs frame-level feature extraction on each of the first convolution processing results to obtain the second frame-level feature sequence of each modality.

[0084] Here, in step d3, each of the first convolution processing results is obtained by performing convolution processing on the silhouette image sequence of each modality through the first convolution module in the first feature extractor, that is, as shown in the example. Figure 2 As shown, the input data of the second feature extractor 202 is each of the first convolution processing results obtained by the first convolution module 210 in the first feature extractor 201 after performing convolution processing on the silhouette image sequence of each modality.

[0085] Specifically, the structure of the second feature extractor 202 is the same as that of the first feature extractor 201. The method by which the second feature extractor 202 performs frame-level feature extraction for each of the first convolution processing results can refer to the method by which the first feature extractor 201 performs frame-level feature extraction for each modality of silhouette image sequence in step d1 above. The repetition will not be repeated here.

[0086] Step d4: Through the ensemble pooling layer, the second frame-level feature sequences of each modality are aggregated to obtain the second ensemble-level features of each modality.

[0087] Here, the specific implementation of step d4 is the same as that of step d2 above, and the repetitions will not be repeated here.

[0088] Step d5: Perform horizontal pyramid pooling on the first set-level features and the second set-level features of each modality to obtain multiple one-dimensional features corresponding to each modality.

[0089] Here, in the feature mapping stage, such as Figure 2 As shown, the gait recognition model can divide the first and second set-level features of each modality into multiple one-dimensional features 210 by performing horizontal pyramid pooling on the first and second set-level features of each modality. The multiple one-dimensional features 210 of the same modality constitute... Figure 2 The feature shown in Figure 203.

[0090] It should be noted that the embodiments of this application do not impose any limitation on the specific number of one-dimensional features corresponding to each mode.

[0091] Step d6: For each modality, the multiple one-dimensional features are mapped to the same discriminant space using an independent fully connected layer corresponding to each one-dimensional feature, thereby obtaining the gait features of each modality for the pedestrian.

[0092] Here, each one-dimensional feature 210 corresponds to an independent fully connected layer ( Figure 2 (Not shown in the image) For each one-dimensional feature, the gait recognition model can use a corresponding independent fully connected layer to map it to the same discriminant space. In this way, through feature mapping, all one-dimensional features of different modalities of the same pedestrian can be mapped to the same discriminant space, so that gait features of different modalities can be spliced ​​together.

[0093] Step d7: Perform gait feature concatenation on the gait features of each modality corresponding to the pedestrian, and use the concatenation result as the gait feature recognition result of the pedestrian.

[0094] It should be noted that when performing the splicing process, it is only necessary to ensure that the different gait features of each modality of the same pedestrian are spliced ​​in series (i.e., connected end to end). This application does not limit the specific modality type corresponding to the two adjacent gait features during splicing.

[0095] Regarding the specific training method of the above gait recognition model, as an optional embodiment, Figure 3 This document illustrates a flowchart of a training method for a gait recognition model provided in an embodiment of this application. Figure 3 As shown, when training the initial model of the gait recognition model, the training method includes steps S301-S303; specifically:

[0096] S301, for each sample object, obtain a sequence of silhouette images of multiple different modalities corresponding to that sample object as a set of training samples.

[0097] Here, each sample object is used to represent a pedestrian. The silhouette image sequence of multiple different modalities corresponding to the same sample object is used as a set of training samples. The specific acquisition method of each set of training samples is the same as the method of acquiring the silhouette image sequence of the same pedestrian in steps S101-S102 above. The repetition will not be repeated here.

[0098] S302, input each group of training samples into the initial model, and output the gait feature prediction result for each sample object.

[0099] Here, the initial model described above represents the gait recognition model before training is complete. The model structure of the initial model can be found in [reference needed]. Figure 2 The gait recognition model shown here will not be repeated here.

[0100] Specifically, the method of outputting the gait feature prediction result in step S302 is the same as the method of outputting the gait feature recognition result in step S103 above, and the repetition will not be repeated here.

[0101] S303, for the gait feature prediction results of each sample object, the gait feature prediction results of the same sample object are used as positive samples, and the gait feature prediction results of different sample objects are used as negative samples. The model parameters of the initial model are adjusted according to the model optimization objective to obtain the initial model including the adjusted model parameters as the gait recognition model.

[0102] Here, the optimization objective of the above model represents that the gait feature prediction result of the sample object is close to the positive sample and far away from the negative sample.

[0103] Specifically, during the initial model training, for each sample object's gait feature prediction result, the first feature distance between the gait feature prediction result of that sample object and the positive sample (i.e., the gait feature prediction result of the same sample object) and the second feature distance between the gait feature prediction result of that sample object and the negative sample (i.e., the gait feature prediction result of different sample objects) can be calculated. The calculated first and second feature distances are substituted into the model loss function to calculate the model loss. By narrowing the first feature distance with the positive sample and widening the second feature distance with the negative sample (i.e., the above-mentioned model optimization objective), the model loss is adjusted (which is also equivalent to adjusting the model parameters) until the model converges, and the initial model including the adjusted model parameters (i.e., the converged initial model) is obtained as the gait recognition model.

[0104] It should be noted that when adjusting the model parameters of the initial model according to the above-mentioned model optimization objective, the specific model loss functions that can be used include, but are not limited to, triplet loss and circle loss; this application embodiment does not limit the specific model loss function used.

[0105] Based on the gait recognition method described in steps S101-S103 above, and for the gait feature recognition results of each pedestrian output by the gait recognition model, this application embodiment also provides a gait feature retrieval method, such as... Figure 4 As shown, Figure 4 This illustration shows a flowchart of a gait feature retrieval method provided in an embodiment of this application. The method includes steps S401-S403; specifically:

[0106] S401, for each pedestrian's gait feature recognition result output by the gait recognition model, the pedestrian's gait feature recognition result is used as the retrieval object, and the similarity between the retrieval object and each gait feature stored in the gait feature database is calculated to obtain the similarity calculation result between the retrieval object and each gait feature.

[0107] Here, each gait feature stored in the gait feature database can be obtained in the manner described in steps S101-S103 above. That is, the specific feature type of each gait feature stored in the gait feature database can refer to the gait feature recognition result obtained in step S103 above.

[0108] It should be noted that, in addition to the above, the gait feature database may also store pedestrian gait features of a single modality. This application embodiment does not limit the specific database content or the specific database construction method of the above gait feature database.

[0109] S402, from the gait feature database, select gait features whose similarity calculation results are greater than or equal to the target threshold as target gait features.

[0110] It should be noted that when performing similarity calculation, the cosine similarity between the search object and each step-state feature can be calculated by matrix multiplication; other similarity calculation methods (such as Euclidean distance) can also be used to calculate the similarity between the search object and each step-state feature. This application embodiment does not limit the specific similarity calculation method used in the above step S401.

[0111] Here, the specific value of the target threshold can be set according to the actual retrieval needs, and this application embodiment does not impose any limitations on this.

[0112] S403, sort each of the selected target gait features according to the similarity calculation results from high to low, and obtain the gait feature retrieval results corresponding to the retrieval object.

[0113] For example, taking a target threshold of 0.7, for the current search object y1, if the target gait features with similarity calculation results greater than or equal to 0.7 are selected from the gait feature database as gait feature r1, gait feature r4, and gait feature r8; where the similarity calculation result between gait feature r1 and search object y1 is 0.8, the similarity calculation result between gait feature r4 and search object y1 is 0.6, and the similarity calculation result between gait feature r8 and search object y1 is 0.7, then, sorting them according to the similarity calculation results from high to low, the gait feature search results corresponding to search object y1 can be obtained as: gait feature r1, gait feature r8, and gait feature r4.

[0114] Based on the same inventive concept, this application also provides a gait recognition device corresponding to the above-mentioned gait recognition method. Since the principle of the gait recognition device in this application is similar to that of the above-mentioned gait recognition method in this application, the implementation of the gait recognition device can refer to the implementation of the above-mentioned gait recognition method, and the repeated parts will not be described again.

[0115] Reference Figure 5 As shown, Figure 5 This illustration shows a structural schematic diagram of a multimodal gait recognition device provided in an embodiment of this application. The gait recognition device includes:

[0116] The data acquisition module 501 is used to simultaneously acquire walking video images of different pedestrians within the same spatial range using multiple image acquisition devices of different modalities, and obtain walking video images of each modality.

[0117] The image processing module 502 is used to perform image target detection on pedestrians included in each mode of walking video images, and to separate the silhouette image sequence belonging to the same pedestrian from each mode of walking video images based on the image target detection results.

[0118] The feature recognition module 503 is used to input a sequence of silhouette images of multiple different modalities corresponding to the same pedestrian into a pre-trained gait recognition model and output a gait feature recognition result for the pedestrian; wherein, the gait feature recognition result represents the splicing result of the gait features of multiple different modalities corresponding to the pedestrian.

[0119] In an optional implementation, when separating the silhouette image sequence belonging to the same pedestrian from the walking video images of each modality based on the image target detection results, the image processing module 502 is configured to:

[0120] For each mode of walking video image, the image target detection results of the walking video image of that mode are segmented into instances to obtain silhouette images of different pedestrians corresponding to that mode and pedestrian labels corresponding to each frame of the silhouette image;

[0121] Based on the pedestrian tag corresponding to each silhouette image, multiple silhouette images corresponding to the same pedestrian tag are selected from the silhouette images of different pedestrians corresponding to this modality.

[0122] In an optional implementation, when the sequence of silhouette images of multiple different modalities corresponding to the same pedestrian is input into a pre-trained gait recognition model, and the gait feature recognition result for that pedestrian is output, the feature recognition module 503 is used to:

[0123] The silhouette image sequences of multiple different modalities corresponding to the same pedestrian are simultaneously input into the same channel of the gait recognition model. The gait recognition model performs mixed feature extraction on the silhouette image sequences of multiple different modalities in the channel and outputs the gait feature recognition result for the pedestrian.

[0124] In an optional implementation, when the sequence of silhouette images of multiple different modalities corresponding to the same pedestrian is input into a pre-trained gait recognition model, and the gait feature recognition result for that pedestrian is output, the feature recognition module 503 is used to:

[0125] The silhouette image sequences of the same pedestrian corresponding to the various different modalities are input into different channels of the gait recognition model. The gait recognition model performs independent feature extraction on the silhouette image sequence in each channel to obtain the gait features of each modality corresponding to the pedestrian.

[0126] The gait features corresponding to each modality of the pedestrian are concatenated, and the concatenated result is output as the gait feature recognition result for the pedestrian.

[0127] In an optional implementation, when the sequence of silhouette images of multiple different modalities corresponding to the same pedestrian is input into a pre-trained gait recognition model, and the gait feature recognition result for that pedestrian is output, the feature recognition module 503 is used to:

[0128] For the silhouette image sequences of the same pedestrian corresponding to the various different modalities, the first feature extractor in the gait recognition model performs frame-level feature extraction on the silhouette image sequence of each modality to obtain the first frame-level feature sequence of each modality;

[0129] The first frame-level feature sequence of each modality is aggregated through the ensemble pooling layer in the gait recognition model to obtain the first ensemble-level feature of each modality;

[0130] For the first convolution processing result of the silhouette image sequence for each modality, the second feature extractor in the gait recognition model performs frame-level feature extraction on each first convolution processing result to obtain the second frame-level feature sequence for each modality; wherein, each first convolution processing result is obtained by performing convolution processing on the silhouette image sequence for each modality through the first convolution module in the first feature extractor.

[0131] The second frame-level feature sequence of each modality is aggregated through the pooling layer to obtain the second pooling-level feature of each modality;

[0132] Horizontal pyramid pooling is performed on the first set-level features and the second set-level features for each modality to obtain multiple one-dimensional features corresponding to each modality;

[0133] For each modality corresponding to the multiple one-dimensional features, each one-dimensional feature corresponding to each modality is mapped to the same discriminant space using an independent fully connected layer corresponding to each one-dimensional feature, thereby obtaining the gait features of each modality corresponding to the pedestrian.

[0134] The gait features corresponding to each modality of the pedestrian are concatenated, and the concatenation result is used as the gait feature recognition result of the pedestrian.

[0135] In an optional implementation, the gait recognition device further includes a model training module, wherein the model training module is used to train the gait recognition model using the following method:

[0136] For each sample object, a sequence of silhouette images in multiple different modalities corresponding to that sample object is obtained as a set of training samples;

[0137] Each set of training samples is input into the initial model, and the gait feature prediction result for each sample object is output.

[0138] For each of the gait feature prediction results of the sample object, the gait feature prediction results of the same sample object are used as positive samples, and the gait feature prediction results of different sample objects are used as negative samples. The model parameters of the initial model are adjusted according to the model optimization objective to obtain the initial model including the adjusted model parameters as the gait recognition model; wherein, the model optimization objective represents that the gait feature prediction result of the sample object is close to the positive sample and far away from the negative sample.

[0139] In an optional embodiment, the gait recognition device further includes a gait retrieval module, wherein the gait retrieval module is used for:

[0140] For each pedestrian's gait feature recognition result output by the gait recognition model, the pedestrian's gait feature recognition result is used as the retrieval object. The similarity between the retrieval object and each gait feature stored in the gait feature database is calculated to obtain the similarity calculation result between the retrieval object and each gait feature.

[0141] From the gait feature database, gait features whose similarity calculation results are greater than or equal to the target threshold are selected as target gait features;

[0142] According to the similarity calculation results from high to low, each of the selected target gait features is sorted to obtain the gait feature retrieval results corresponding to the retrieval object.

[0143] like Figure 6 As shown, this application provides a computer device 600 for executing the multimodal gait recognition method of this application. The device includes a memory 601, a processor 602, and a computer program stored in the memory 601 and executable on the processor 602. The memory 601 and the processor 602 communicate via a bus. When the processor 602 executes the computer program, it implements the steps of the multimodal gait recognition method described above.

[0144] Specifically, the memory 601 and processor 602 mentioned above can be general-purpose memory and processor, without any specific limitations. When the processor 602 runs the computer program stored in the memory 601, it can execute the above-mentioned multimodal gait recognition method.

[0145] Corresponding to the multimodal gait recognition method in this application, this application embodiment also provides a computer-readable storage medium storing a computer program, which is executed by a processor to perform the steps of the multimodal gait recognition method described above.

[0146] Specifically, the storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the storage medium is run, it can execute the above-mentioned multimodal gait recognition method.

[0147] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. The system embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and there may be other division methods in actual implementation. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between systems or units may be electrical, mechanical, or other forms.

[0148] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0149] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0150] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0151] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0152] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A multi-modal based gait recognition method, characterized in that, The gait recognition method comprises: In the same spatial range, a plurality of different modal image acquisition devices are used to simultaneously collect walking video images of different pedestrians in the spatial range, to obtain walking video images of each modal; Image target detection is performed on pedestrians included in the walking video images of each modal, and according to the image target detection result, a silhouette image sequence belonging to the same pedestrian is separated from the walking video images of each modal; The silhouette image sequences of the same pedestrian corresponding to a plurality of different modalities are input into a pre-trained gait recognition model, and a gait feature recognition result for the pedestrian is output; wherein the gait feature recognition result represents the splicing result of the gait features of the same pedestrian corresponding to a plurality of different modalities. The gait recognition method further comprises: For each gait feature recognition result of the pedestrian output by the gait recognition model, the gait feature recognition result of the pedestrian is taken as a search object, the similarity between the search object and each gait feature stored in a gait feature database is calculated, and a similarity calculation result between the search object and each gait feature is obtained; From the gait feature database, a gait feature with a similarity calculation result greater than or equal to a target threshold is selected as a target gait feature; According to the order from high to low of the similarity calculation result, each target gait feature selected is sorted, and a gait feature search result corresponding to the search object is obtained.

2. The gait recognition method according to claim 1, characterized in that, According to the image target detection result, a silhouette image sequence belonging to the same pedestrian is separated from the walking video images of each modal, comprising: For each modal walking video image, instance segmentation is performed on the image target detection result of the modal walking video image, to obtain a silhouette image of different pedestrians corresponding to the modal and a pedestrian label corresponding to each frame of the silhouette image; According to the pedestrian label corresponding to each frame of the silhouette image, a plurality of frames of silhouette images corresponding to the same pedestrian label are selected from the silhouette images of different pedestrians corresponding to the modal.

3. The gait recognition method of claim 1, wherein, The silhouette image sequences of the same pedestrian corresponding to a plurality of different modalities are input into a pre-trained gait recognition model, and a gait feature recognition result for the pedestrian is output, comprising: The silhouette image sequences of the same pedestrian corresponding to the plurality of different modalities are simultaneously input into the same channel of the gait recognition model, mixed feature extraction is performed on the silhouette image sequences of the plurality of different modalities in the channel by the gait recognition model, and the gait feature recognition result for the pedestrian is output.

4. The gait recognition method of claim 1, wherein, The silhouette image sequences of the same pedestrian corresponding to the plurality of different modalities are input into a pre-trained gait recognition model, and a gait feature recognition result for the pedestrian is output, further comprising: The silhouette image sequences of the same pedestrian corresponding to the plurality of different modalities are respectively input into different channels of the gait recognition model, independent feature extraction is performed on the silhouette image sequences in each channel by the gait recognition model, and the gait feature of each modal corresponding to the pedestrian is obtained; The gait features of each modality corresponding to the pedestrian are spliced, and the splicing result is output as the gait feature recognition result of the pedestrian.

5. The gait recognition method of claim 1, wherein, The gait feature recognition result of the pedestrian is output by inputting the silhouette image sequences of the same pedestrian in different modalities into the pre-trained gait recognition model, including: For the silhouette image sequences of the same pedestrian in different modalities, the first frame-level feature sequence of each modality is obtained by performing frame-level feature extraction on the silhouette image sequence of each modality through the first feature extractor in the gait recognition model. The first set-level feature of each modality is obtained by aggregating the first frame-level feature sequence of each modality through the set pooling layer in the gait recognition model. For the first convolution processing result of the silhouette image sequence of each modality, the second frame-level feature sequence of each modality is obtained by performing frame-level feature extraction on each first convolution processing result through the second feature extractor in the gait recognition model; wherein each first convolution processing result is obtained by performing convolution processing on the silhouette image sequence of each modality through the first convolution module in the first feature extractor. The second set-level feature of each modality is obtained by aggregating the second frame-level feature sequence of each modality through the set pooling layer. The first set-level feature and the second set-level feature of each modality are respectively subjected to horizontal pyramid pooling to obtain a plurality of one-dimensional features corresponding to each modality. For the plurality of one-dimensional features corresponding to each modality, each one-dimensional feature corresponding to each modality is mapped to the same discriminant space by using an independent fully connected layer corresponding to each one-dimensional feature, to obtain the gait feature of each modality corresponding to the pedestrian. The gait features of each modality corresponding to the pedestrian are spliced, and the splicing result is output as the gait feature recognition result of the pedestrian.

6. The gait recognition method of claim 1, wherein, The gait recognition model is trained by the following method: For each sample object, a plurality of silhouette image sequences of different modalities corresponding to the sample object are obtained as a group of training samples; Each group of training samples is input into an initial model, and the gait feature prediction result of each sample object is output; For the gait feature prediction result of each sample object, the gait feature prediction result of the same sample object is taken as a positive sample, the gait feature prediction result of different sample objects is taken as a negative sample, and the model parameters of the initial model are adjusted according to the model optimization target to obtain the initial model including the adjusted model parameters as the gait recognition model; wherein the model optimization target represents that the gait feature prediction result of the sample object is close to the positive sample and far away from the negative sample. 7.A multi-modal based gait recognition device, characterized in that, The gait recognition device includes: A data acquisition module is configured to simultaneously acquire walking video images of different pedestrians in a same space range by using image acquisition devices of different modalities, to obtain walking video images of each modality. An image processing module is configured to perform image target detection on pedestrians included in the walking video images of each modality, and separate silhouette image sequences of the same pedestrian from the walking video images of each modality according to the image target detection results. A feature recognition module is configured to input the silhouette image sequences of the same pedestrian in different modalities into a pre-trained gait recognition model, and output a gait feature recognition result of the pedestrian. The gait recognition device further comprises a gait retrieval module, wherein the gait retrieval module is configured to: For each gait feature recognition result of the pedestrian output by the gait recognition model, take the gait feature recognition result of the pedestrian as a retrieval object, calculate the similarity between the retrieval object and each gait feature stored in a gait feature database, and obtain a similarity calculation result between the retrieval object and each gait feature. From the gait feature database, filter out gait features with a similarity calculation result greater than or equal to a target threshold as target gait features. Sort each target gait feature filtered out according to the similarity calculation result from high to low, and obtain a gait feature retrieval result corresponding to the retrieval object.

8. An electronic device, comprising: It comprises: A processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor and the memory communicate through the bus, the machine readable instructions are executed by the processor to execute the steps of the multi-modal based gait recognition method in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which is executed by the processor to execute the steps of the multi-modal based gait recognition method in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Gait visual angle detection method and device and gait recognition method and device

    CN114445910A