Methods, devices, electronic equipment and media for generating gait recognition training data
By merging the posture and appearance features of sample pedestrians in the gait recognition model, a new gait sequence is generated, which solves the problem of decreased recognition accuracy caused by limited appearance variations and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202310547729.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-16
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2043-05-16
AI Technical Summary
Existing gait recognition models are unable to learn robust overall posture features to appearance changes due to the limited variation in the appearance of pedestrians in the training data, which affects recognition accuracy.
By inputting the gait sequences of sample pedestrians into a pre-trained encoder to extract pose and appearance features, and merging them to generate new gait sequences, the training sample set is expanded, increasing the diversity of appearance variations.
Training samples with rich appearances were generated, which improved the recognition accuracy of the gait recognition model.
Smart Images

Figure CN116486487B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gait recognition technology, and in particular to a method, apparatus, electronic device, and medium for generating gait recognition training data. Background Technology
[0002] Gait recognition is one of the most valuable long-range biometric technologies. However, changes in appearance due to clothing, footwear, and accessories pose the biggest bottleneck to gait recognition. Specifically, in the training data of gait recognition models, because the appearance variations of each sample pedestrian are limited (i.e., the clothing, footwear, and accessories of the same sample pedestrian are usually unchanged), the gait recognition model cannot learn overall posture features that are robust to appearance changes, thus leading to a decline in recognition performance.
[0003] For example, when a gait recognition model trained on training data with limited appearance variables is applied, after obtaining the target gait sequence of a pedestrian, it typically extracts the target gait features from the sequence. These features are then compared with gait features in a database to determine if the pedestrian matches the database entry. However, if the target pedestrian is indeed in the database, significant differences in clothing between the target and database entries can lead to a large discrepancy between the extracted gait features and those in the database. This results in a low similarity between the two gait features, affecting the accuracy of the identification and leading to inaccurate results. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide a method, apparatus, electronic device and medium for generating gait recognition training data, so as to increase the appearance variation of pedestrians in the model training samples, generate model training samples with rich appearance, expand the model training sample set, and thus improve the recognition accuracy of the gait recognition model.
[0005] In a first aspect, embodiments of this application provide a method for generating gait recognition training data, including:
[0006] The first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian are input into the pre-trained encoder to obtain the first pose feature and the first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and the second pose feature and the second appearance feature of the second sample pedestrian corresponding to each second step state image frame in the second step state sequence.
[0007] For each of the first appearance features, the first appearance feature is merged with each of the second posture features to obtain a plurality of first merged features corresponding to the first appearance feature; and for each of the second appearance features, the second appearance feature is merged with each of the first posture features to obtain a plurality of second merged features corresponding to the second appearance feature.
[0008] Multiple first merged features corresponding to the same first appearance feature are input into a pre-trained decoder to generate a third gait sequence of the second sample pedestrian with the same first appearance feature; and multiple second merged features corresponding to the same second appearance feature are input into the decoder to generate a fourth gait sequence of the first sample pedestrian with the same second appearance feature.
[0009] In conjunction with the first aspect, embodiments of this application provide a first possible implementation of the first aspect, wherein, before inputting the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, the method further includes:
[0010] Obtain the first gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian from the original training set;
[0011] After generating the third gait sequence and the fourth gait sequence, the method further includes:
[0012] The third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian are added to the original training set to obtain a new training set; the new training set contains multiple gait sequences corresponding to each sample pedestrian, and each gait sequence corresponding to the same sample pedestrian contains its own appearance.
[0013] In conjunction with the first aspect, this application provides a second possible implementation of the first aspect, wherein, before inputting the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, the method further includes:
[0014] The fifth and sixth gait sequences of the third sample pedestrian under different appearances are input into the initial encoder to be trained to obtain the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence.
[0015] For each of the third posture features, the third posture feature is merged with each of the third appearance features corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the fourth appearance features corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature.
[0016] For each of the third pose features corresponding to each of the third merged features, the third merged feature is converted into a first synthetic image frame of the third sample pedestrian by the initial decoder to be trained, so as to obtain the first synthetic image frame corresponding to each of the third merged features; and for each of the fourth pose features corresponding to each of the fourth merged features, the fourth merged feature is converted into a second synthetic image frame of the third sample pedestrian by the initial decoder.
[0017] For each first synthesized image frame corresponding to each of the third pose features, calculate the first pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthesized image frame at each pixel point to obtain the first pixel average of each first synthesized image frame; and for each second synthesized image frame corresponding to each of the fourth pose features, calculate the second pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthesized image frame at each pixel point to obtain the second pixel average of each second synthesized image frame.
[0018] The average of all first pixel values corresponding to all third pose features and the average of all second pixel values corresponding to all fourth pose features are calculated to obtain the third pixel average value corresponding to the third sample pedestrian, and the third pixel average value is used as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0019] In conjunction with the second possible implementation of the first aspect, this application provides a third possible implementation of the first aspect, wherein, after obtaining the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence, the method further includes:
[0020] For each of the third posture features, the third posture feature is merged with each of the fourth appearance features to obtain a plurality of fifth merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the third appearance features to obtain a plurality of sixth merged features corresponding to the fourth posture feature.
[0021] For each fifth merged feature corresponding to each third pose feature, the fifth merged feature is converted into a third synthetic image frame of the third sample pedestrian by the initial decoder; and for each sixth merged feature corresponding to each fourth pose feature, the sixth merged feature is converted into a fourth synthetic image frame of the third sample pedestrian by the initial decoder.
[0022] For each of the third synthetic image frames corresponding to each of the third pose features, calculate the fourth pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the third synthetic image frame at each pixel point, to obtain the fourth pixel average of each of the third synthetic image frames; and for each of the fourth synthetic image frames corresponding to each of the fourth pose features, calculate the fifth pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the fourth synthetic image frame at each pixel point, to obtain the fifth pixel average of each of the fourth synthetic image frames;
[0023] The average of the average of all fourth pixels corresponding to all third pose features and the average of all fifth pixels corresponding to all fourth pose features are calculated to obtain the average of the sixth pixel corresponding to the third sample pedestrian, and the average of the sixth pixel is used as the second loss value; the second loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0024] In conjunction with the third possible implementation of the first aspect, this application provides a fourth possible implementation of the first aspect, wherein the method further includes:
[0025] The seventh gait sequence of the fourth sample pedestrian is input into the initial encoder to obtain the fifth appearance feature and fifth pose feature corresponding to each seventh gait image frame in the seventh gait sequence;
[0026] For each of the third posture features, the third posture feature is merged with each of the fifth appearance features to obtain a plurality of seventh merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the fifth appearance features to obtain a plurality of eighth merged features corresponding to the fourth posture feature.
[0027] For each of the seventh merged features corresponding to each of the third pose features, the seventh merged feature is converted into a fifth composite image frame of the third sample pedestrian by the initial decoder; and for each of the eighth merged features corresponding to each of the fourth pose features, the eighth merged feature is converted into a sixth composite image frame of the third sample pedestrian by the initial decoder.
[0028] Each of the fifth synthesized image frames is input into the initial encoder to obtain the sixth pose feature and the sixth appearance feature corresponding to each of the fifth synthesized image frames; and each of the sixth synthesized image frames is input into the initial encoder to obtain the seventh pose feature and the seventh appearance feature corresponding to each of the sixth synthesized image frames.
[0029] For each of the sixth posture features, the sixth posture feature is merged with each of the third appearance features to obtain multiple ninth merged features corresponding to the sixth posture feature; and for each of the seventh posture features, the seventh posture feature is merged with each of the third appearance features to obtain multiple tenth merged features corresponding to the seventh posture feature.
[0030] For each of the ninth merged features corresponding to the sixth pose feature, the initial decoder converts the ninth merged feature into the seventh composite image frame of the third sample pedestrian; and for each of the tenth merged features corresponding to the seventh pose feature, the initial decoder converts the tenth merged feature into the eighth composite image frame of the third sample pedestrian.
[0031] For each of the seventh synthesized image frames corresponding to each of the sixth pose features, calculate the seventh pixel average of the pixel difference between the third gait image frame corresponding to the sixth pose feature and the seventh synthesized image frame at each pixel point; and for each of the eighth synthesized image frames corresponding to each of the seventh pose features, calculate the eighth pixel average of the pixel difference between the fourth gait image frame corresponding to the seventh pose feature and the eighth synthesized image frame at each pixel point.
[0032] The average of all seventh pixel values corresponding to all sixth pose features and the average of all eighth pixel values corresponding to all seventh pose features are calculated to obtain the ninth pixel average value, which is used as the third loss value. The third loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0033] In conjunction with the third possible implementation of the first aspect, this application provides a fifth possible implementation of the first aspect, wherein the method further includes:
[0034] Calculate the average feature of each of the third pose features of the third sample pedestrian to obtain the first pose average feature; and calculate the average feature of each of the fourth pose features of the third sample pedestrian to obtain the second pose average feature;
[0035] The difference between the average features of the first pose and the average features of the second pose is calculated to obtain a fourth loss value; the fourth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0036] In conjunction with the fourth possible implementation of the first aspect, this application provides a sixth possible implementation of the first aspect, wherein the method further includes:
[0037] For each of the third pose features of the third sample pedestrian, the third pose feature is merged with each of the fifth appearance features to obtain multiple eleventh merged features corresponding to the third pose feature;
[0038] Using the InfoNCE Loss function, a first InfoNCE loss value is calculated using the third pose feature, the third appearance feature, and the eleventh merged feature; and using the InfoNCE Loss function, a second InfoNCE loss value is calculated using the fifth appearance feature, the fifth pose feature, and the eleventh merged feature.
[0039] The sum of the first InfoNCE loss value and the second InfoNCE loss value is determined as the fifth loss value; the fifth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0040] Secondly, embodiments of this application also provide an apparatus for generating gait recognition training data, comprising:
[0041] The first input module is used to input the first step state sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder to obtain the first pose feature and the first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and to obtain the second pose feature and the second appearance feature of the second sample pedestrian corresponding to each second gait image frame in the second gait sequence.
[0042] The first merging module is configured to merge each first appearance feature with each second posture feature to obtain a plurality of first merged features corresponding to the first appearance feature; and to merge each second appearance feature with each first posture feature to obtain a plurality of second merged features corresponding to the second appearance feature.
[0043] The generation module is configured to input multiple first merged features corresponding to the same first appearance feature into a pre-trained decoder to generate a third gait sequence of the second sample pedestrian with the same first appearance feature; and to input multiple second merged features corresponding to the same second appearance feature into the decoder to generate a fourth gait sequence of the first sample pedestrian with the same second appearance feature.
[0044] In conjunction with the second aspect, embodiments of this application provide a first possible implementation of the second aspect, which further includes:
[0045] The acquisition module is used to acquire the first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian from the original training set before the first input module inputs the first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian into the pre-trained encoder;
[0046] The addition module is used to add the third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian to the original training set after the generation module generates the third gait sequence and the fourth gait sequence, to obtain a new training set; the new training set contains multiple gait sequences corresponding to each sample pedestrian, and each gait sequence corresponding to the same sample pedestrian contains its own appearance.
[0047] In conjunction with the second aspect, embodiments of this application provide a second possible implementation of the second aspect, which further includes:
[0048] The second input module is used to input the fifth and sixth gait sequences of the third sample pedestrian under different appearances into the initial encoder to be trained before the first input module inputs the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, so as to obtain the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence.
[0049] The second merging module is used to merge each of the third posture features with each of the third appearance features corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and to merge each of the fourth posture features with each of the fourth appearance features corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature.
[0050] The first conversion module is used to convert each of the third merged features corresponding to each of the third pose features into a first synthetic image frame of the third sample pedestrian through an initial decoder to be trained, thereby obtaining a first synthetic image frame corresponding to each of the third merged features; and to convert each of the fourth merged features corresponding to each of the fourth pose features into a second synthetic image frame of the third sample pedestrian through the initial decoder.
[0051] The first calculation module is configured to, for each of the first synthetic image frames corresponding to the third pose feature, calculate the first average pixel value of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthetic image frame at each pixel point, to obtain the first average pixel value corresponding to each of the first synthetic image frames; and for each of the second synthetic image frames corresponding to the fourth pose feature, calculate the second average pixel value of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthetic image frame at each pixel point, to obtain the second average pixel value corresponding to each of the second synthetic image frames.
[0052] The second calculation module is used to calculate the average of all the first pixel averages corresponding to all the third pose features and the average of all the second pixel averages corresponding to all the fourth pose features, to obtain the third pixel average corresponding to the third sample pedestrian, and to use the third pixel average as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0053] In conjunction with the second possible implementation of the second aspect, this application provides a third possible implementation of the second aspect, which further includes:
[0054] The third merging module is used to, after obtaining the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence from the second input module, merge the third pose feature with each of the fourth appearance features to obtain multiple fifth merged features corresponding to the third pose feature; and merge the fourth pose feature with each of the third appearance features to obtain multiple sixth merged features corresponding to the fourth pose feature.
[0055] The second conversion module is used to convert each fifth merged feature corresponding to each third pose feature into a third synthetic image frame of the third sample pedestrian through the initial decoder; and to convert each sixth merged feature corresponding to each fourth pose feature into a fourth synthetic image frame of the third sample pedestrian through the initial decoder.
[0056] The third calculation module is used to calculate, for each of the third synthetic image frames corresponding to the third pose feature, the fourth pixel average value of the pixel difference between the third gait image frame corresponding to the third pose feature and the third synthetic image frame at each pixel point, to obtain the fourth pixel average value corresponding to each of the third synthetic image frames; and to calculate, for each of the fourth synthetic image frames corresponding to the fourth pose feature, the fifth pixel average value of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the fourth synthetic image frame at each pixel point, to obtain the fifth pixel average value corresponding to each of the fourth synthetic image frames.
[0057] The fourth calculation module is used to calculate the average of all the fourth pixel averages corresponding to all the third pose features and the average of all the fifth pixel averages corresponding to all the fourth pose features, to obtain the sixth pixel average corresponding to the third sample pedestrian, and to use the sixth pixel average as the second loss value; the second loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0058] In conjunction with the third possible implementation of the second aspect, this application provides a fourth possible implementation of the second aspect, which further includes:
[0059] The third input module is used to input the seventh gait sequence of the fourth sample pedestrian into the initial encoder to obtain the fifth appearance feature and fifth pose feature corresponding to each seventh gait image frame in the seventh gait sequence;
[0060] The fourth merging module is used to merge each of the third posture features with each of the fifth appearance features to obtain a plurality of seventh merged features corresponding to the third posture feature; and to merge each of the fourth posture features with each of the fifth appearance features to obtain a plurality of eighth merged features corresponding to the fourth posture feature.
[0061] The third conversion module is used to convert each of the seventh merged features corresponding to each of the third pose features into a fifth synthetic image frame of the third sample pedestrian through the initial decoder; and to convert each of the eighth merged features corresponding to each of the fourth pose features into a sixth synthetic image frame of the third sample pedestrian through the initial decoder.
[0062] The fourth input module is used to input each of the fifth synthesized image frames into the initial encoder to obtain the sixth pose feature and the sixth appearance feature corresponding to each of the fifth synthesized image frames; and to input each of the sixth synthesized image frames into the initial encoder to obtain the seventh pose feature and the seventh appearance feature corresponding to each of the sixth synthesized image frames.
[0063] The fifth merging module is used to merge each of the sixth posture features with each of the third appearance features to obtain multiple ninth merged features corresponding to the sixth posture feature; and to merge each of the seventh posture features with each of the third appearance features to obtain multiple tenth merged features corresponding to the seventh posture feature.
[0064] The fourth conversion module is used to convert each ninth merged feature corresponding to each sixth pose feature into a seventh composite image frame of the third sample pedestrian through the initial decoder; and to convert each tenth merged feature corresponding to each seventh pose feature into an eighth composite image frame of the third sample pedestrian through the initial decoder.
[0065] The fifth calculation module is used to calculate, for each of the seventh composite image frames corresponding to the sixth pose feature, the seventh pixel average of the pixel difference between the third gait image frame corresponding to the sixth pose feature and the seventh composite image frame at each pixel point; and to calculate, for each of the eighth composite image frames corresponding to the seventh pose feature, the eighth pixel average of the pixel difference between the fourth gait image frame corresponding to the seventh pose feature and the eighth composite image frame at each pixel point.
[0066] The sixth calculation module is used to calculate the average of the average values of all seventh pixels corresponding to all sixth pose features and the average values of all eighth pixels corresponding to all seventh pose features to obtain the average value of the ninth pixel, and use the average value of the ninth pixel as the third loss value; the third loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0067] In conjunction with the third possible implementation of the second aspect, this application provides a fifth possible implementation of the second aspect, which further includes:
[0068] The seventh calculation module is used to calculate the average feature of each of the third posture features of the third sample pedestrian to obtain the first posture average feature; and to calculate the average feature of each of the fourth posture features of the third sample pedestrian to obtain the second posture average feature.
[0069] The eighth calculation module is used to calculate the difference between the average features of the first pose and the average features of the second pose to obtain a fourth loss value; the fourth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0070] In conjunction with the fourth possible implementation of the second aspect, this application provides a sixth possible implementation of the second aspect, which further includes:
[0071] The sixth merging module is used to merge the third pose feature with each of the fifth appearance features for each of the third pose features of the third sample pedestrian to obtain multiple eleventh merged features corresponding to the third pose feature.
[0072] The ninth calculation module is used to calculate a first InfoNCE loss value using the third pose feature, the third appearance feature, and the eleventh merged feature through the InfoNCE Loss function; and to calculate a second InfoNCE loss value using the fifth appearance feature, the fifth pose feature, and the eleventh merged feature through the InfoNCE Loss function.
[0073] A determination module is used to determine the sum of the first InfoNCE loss value and the second InfoNCE loss value as a fifth loss value; the fifth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0074] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps in any of the possible implementations of the first aspect described above are performed.
[0075] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps in any of the possible implementations of the first aspect described above.
[0076] This application provides a method, apparatus, electronic device, and medium for generating gait recognition training data. It combines a first appearance feature of a first sample pedestrian with a second posture feature of a second sample pedestrian to obtain a third gait sequence of the second sample pedestrian with the first appearance feature (i.e., a new sample of the second sample pedestrian after changing clothes); and combines the first posture feature of the first sample pedestrian with the second appearance feature of the second sample pedestrian to obtain a fourth gait sequence of the first sample pedestrian with the second appearance feature (i.e., a new sample of the first sample pedestrian after changing clothes). This method increases the appearance variations of the first and second sample pedestrians, which is beneficial for generating model training samples with rich appearances, expanding the model training sample set, and improving the recognition accuracy of the gait recognition model.
[0077] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0078] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0079] Figure 1 A flowchart illustrating a method for generating gait recognition training data according to an embodiment of this application is shown;
[0080] Figure 2 A flowchart is shown below illustrating another method for generating gait recognition training data provided in an embodiment of this application;
[0081] Figure 3 This illustration shows a schematic diagram of a gait recognition training data generation device provided in an embodiment of this application;
[0082] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0083] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0084] Considering that in the training data of gait recognition models, when the appearance variations of sample pedestrians are limited (i.e., the clothing, shoes, hats, and accessories of the same sample pedestrian are usually unchanged), the gait recognition model is prone to failing to learn overall posture features robust to appearance changes, thus leading to a decline in recognition performance. Based on this, embodiments of this application provide a method, apparatus, electronic device, and medium for generating gait recognition training data to increase the appearance variations of sample pedestrians in the model training samples, generate model training samples with rich appearances, expand the model training sample set, and thereby improve the recognition accuracy of the gait recognition model. The following embodiments describe these embodiments.
[0085] Example 1:
[0086] To facilitate understanding of this embodiment, a method for generating gait recognition training data disclosed in this application embodiment will first be described in detail. Figure 1 A flowchart illustrating a method for generating gait recognition training data according to an embodiment of this application is shown, as follows: Figure 1 As shown, the process includes the following steps S101-S103:
[0087] S101: Input the first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian into the pre-trained encoder to obtain the first pose feature and first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and obtain the second pose feature and second appearance feature of the second sample pedestrian corresponding to each second step state image frame in the second step state sequence.
[0088] In this embodiment, the first gait sequence includes multiple consecutive first gait image frames of the first sample pedestrian, and the second gait sequence includes multiple consecutive second gait image frames of the second sample pedestrian.
[0089] The first step state sequence of the first sample pedestrian is input into a pre-trained encoder. The encoder outputs the first pose feature and the first appearance feature of the first sample pedestrian for each first step state image frame in the first step state sequence. The first pose feature is used to characterize the pose information of the first sample pedestrian in the corresponding first step state image frame, and the first appearance feature is used to characterize the appearance information of the first sample pedestrian in the corresponding first step state image frame, such as clothing (clothes, shoes, pants), backpack, hat, etc.
[0090] Similarly, the second gait sequence of the second sample pedestrian is input into a pre-trained encoder. This encoder outputs the second pose feature and the second appearance feature of the second sample pedestrian corresponding to each second gait image frame in the second gait sequence. The second pose feature is used to characterize the pose information of the second sample pedestrian in its corresponding second gait image frame, and the second appearance feature is used to characterize the appearance information of the second sample pedestrian in its corresponding second gait image frame.
[0091] Each first gait sequence corresponds to multiple first pose features and multiple first appearance features, and each second gait sequence corresponds to multiple second pose features and multiple second appearance features.
[0092] In one possible implementation, before performing step S101, specifically: obtaining the first gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian from the original training set.
[0093] In this embodiment, the original training set contains the gait sequence corresponding to each sample pedestrian, wherein the gait sequence is generated from multiple consecutive frames of gait images of the sample pedestrian. The first sample pedestrian is any sample pedestrian in the original training set, and the second sample pedestrian is any sample pedestrian in the original training set other than the first sample pedestrian.
[0094] S102: For each first appearance feature, merge the first appearance feature with each second pose feature to obtain multiple first merged features corresponding to the first appearance feature; and for each second appearance feature, merge the second appearance feature with each first pose feature to obtain multiple second merged features corresponding to the second appearance feature.
[0095] For example, suppose the first gait sequence corresponds to 5 first pose features and 5 first appearance features, and the second gait sequence corresponds to 5 second pose features and 5 second appearance features. Then, taking one of the first appearance features as an example, this first appearance feature is merged with each of the 5 second pose features to obtain 5 first merged features corresponding to that first appearance feature. Thus, there are 25 first merged features from the 5 first appearance features.
[0096] In this embodiment, the first merged feature includes the first appearance feature of the first sample pedestrian and the second pose feature of the second sample pedestrian. The second merged feature includes the second appearance feature of the second sample pedestrian and the first pose feature of the first sample pedestrian.
[0097] S103: Input multiple first merged features corresponding to the same first appearance feature into a pre-trained decoder to generate a third gait sequence of a second sample pedestrian with the same first appearance feature; and input multiple second merged features corresponding to the same second appearance feature into a decoder to generate a fourth gait sequence of a first sample pedestrian with the same second appearance feature.
[0098] In this embodiment, the third gait sequence of the second sample pedestrian and the second gait sequence are two gait sequences with different appearances but the same posture. That is, the appearance information of the second sample pedestrian contained in the second gait sequence is different from the appearance information of the second sample pedestrian contained in the third gait sequence, but the posture information is the same.
[0099] Furthermore, the fourth gait sequence and the first gait sequence of the first sample pedestrian are two gait sequences with different appearances but the same posture.
[0100] In one possible implementation, after generating the third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian, it is further possible to: add the third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian to the original training set to obtain a new training set; the new training set contains multiple gait sequences corresponding to each sample pedestrian, and each gait sequence corresponding to the same sample pedestrian contains its own appearance.
[0101] After obtaining the new training set, the gait recognition model can be trained using each gait sequence of each pedestrian sample included in the new training set. Since each pedestrian sample in the new training set has rich appearance variations, using the new training set to train the gait recognition model is beneficial to improving the recognition accuracy of the target gait recognition model after training.
[0102] In one possible implementation, Figure 2 A flowchart illustrating another method for generating gait recognition training data provided in an embodiment of this application is shown, such as... Figure 2As shown, before performing step S101, the encoder and decoder can be trained using the following steps S1001-S1005:
[0103] S1001: Input the fifth and sixth gait sequences of the third sample pedestrian under different appearances into the initial encoder to be trained, and obtain the third appearance features and third pose features corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance features and fourth pose features corresponding to each fourth gait image frame in the sixth gait sequence.
[0104] In this embodiment, the fifth gait sequence and the sixth gait sequence were collected from the same third sample pedestrian under different appearances, such as different clothing and accessories.
[0105] For example, the fifth gait sequence A1 of the third sample pedestrian A includes three third gait image frames: A11, A12, and A13, and the sixth gait sequence A2 includes three fourth gait image frames: A21, A22, and A23. Then, the fifth gait sequence A1 is input into the initial encoder to be trained. The initial encoder outputs the third appearance feature aw11 and the third pose feature az11 corresponding to the third gait image frame A11, the third appearance feature aw12 and the third pose feature az12 corresponding to the third gait image frame A12, and the third appearance feature aw13 and the third pose feature az13 corresponding to the third gait image frame A13.
[0106] Similarly, the sixth gait sequence A2 is input into the initial encoder, and the fourth appearance feature aw21 and the fourth posture feature az21 corresponding to the fourth gait image frame A21, the fourth appearance feature aw22 and the fourth posture feature az22 corresponding to the fourth gait image frame A22, and the fourth appearance feature aw23 and the fourth posture feature az23 corresponding to the fourth gait image frame A23 are output.
[0107] S1002: For each third posture feature, merge the third posture feature with each third appearance feature corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and for each fourth posture feature, merge the fourth posture feature with each fourth appearance feature corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature.
[0108] For example, the third pose features include az11, az12, and az13. Taking the third pose feature az11 as an example, this third pose feature az11 is merged with the corresponding third appearance features of the third pose features az12 and az13, respectively. That is, the third pose feature az11 is merged with the third appearance features aw12 and aw13, respectively, to obtain two merged third features. In this case, each third pose feature corresponds to two merged third features, and three third pose features correspond to six merged third features.
[0109] Similarly, the fourth pose features include az21, az22, and az23. Taking the fourth pose feature az21 as an example, this fourth pose feature az21 is merged with the corresponding third appearance features of the fourth pose features az22 and az23, respectively. That is, the fourth pose feature az21 is merged with the fourth appearance features aw22 and aw23, respectively, to obtain two fourth merged features. Each fourth pose feature corresponds to two fourth merged features, and three fourth pose features correspond to six fourth merged features.
[0110] S1003: For each third merged feature corresponding to each third pose feature, the third merged feature is converted into the first synthetic image frame of the third sample pedestrian through the initial decoder to be trained, so as to obtain the first synthetic image frame corresponding to each third merged feature; and for each fourth merged feature corresponding to each fourth pose feature, the fourth merged feature is converted into the second synthetic image frame of the third sample pedestrian through the initial decoder.
[0111] For example, taking the third pose feature az11 as an example, each third merged feature corresponding to the third pose feature az11 is input into the initial decoder to be trained. The initial decoder converts the third merged feature into the first synthetic image frame of the third sample pedestrian. Each third merged feature corresponds to one first synthetic image frame, and each third pose feature corresponds to two first synthetic image frames. The two first synthetic image frames corresponding to the third pose feature az11 respectively contain the third pose feature az11 and the third appearance feature aw12, and the first synthetic image frame contains the third pose feature az11 and aw13 respectively.
[0112] Taking the four-pose feature az21 as an example, each fourth merged feature corresponding to the four-pose feature az21 is input into the initial decoder. The initial decoder converts the fourth merged feature into a second synthetic image frame of the third sample pedestrian. Each fourth merged feature corresponds to one second synthetic image frame, and each fourth pose feature corresponds to two fourth merged features; therefore, each fourth pose feature corresponds to two second synthetic image frames. The two second synthetic image frames corresponding to the fourth pose feature az21 respectively contain: the fourth pose feature az21 and the fourth appearance feature aw22, and the second synthetic image frame contains the fourth pose feature az21 and the fourth appearance feature aw23.
[0113] S1004: For each first synthetic image frame corresponding to each third pose feature, calculate the first pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthetic image frame at each pixel point, to obtain the first pixel average corresponding to each first synthetic image frame; and for each second synthetic image frame corresponding to each fourth pose feature, calculate the second pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthetic image frame at each pixel point, to obtain the second pixel average corresponding to each second synthetic image frame.
[0114] For example, taking the third pose feature az11 as an example, this third pose feature az11 corresponds to two first synthetic image frames. Taking one of the first synthetic image frames as an example, the pixel difference between the third gait image frame A11 corresponding to the third pose feature az11 and the first synthetic image frame at each pixel point is calculated to obtain the pixel difference at each pixel point. Then, the average value of all pixel differences is calculated to obtain the first pixel average value. In this embodiment, each third pose feature corresponds to two first synthetic image frames, therefore each third pose feature corresponds to two first pixel average values, and three third pose features correspond to six first pixel average values.
[0115] Similarly, each fourth pose feature corresponds to the average of two second pixels.
[0116] S1005: Calculate the average of all first pixel values corresponding to all third pose features and the average of all second pixel values corresponding to all fourth pose features to obtain the average third pixel value corresponding to the third sample pedestrian, and use the average third pixel value as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0117] In this embodiment, the learnable parameters in the initial encoder and the learnable parameters in the initial decoder are trained using a first loss value until the first loss value converges, at which point training stops, resulting in the trained encoder and decoder. The smaller the first loss value, the better the initial encoder can distinguish between pose features and appearance features in gait features, and the better the initial decoder can convert the merged features into synthetic image frames.
[0118] In one possible implementation, after step S1001 is completed, the following steps S1006-S1009 can also be performed:
[0119] S1006: For each third pose feature, merge the third pose feature with each fourth appearance feature to obtain multiple fifth merged features corresponding to the third pose feature; and for each fourth pose feature, merge the fourth pose feature with each third appearance feature to obtain multiple sixth merged features corresponding to the fourth pose feature.
[0120] For example, the third pose features include az11, az12, and az13. Taking the third pose feature az11 as an example, the third pose feature az11 is merged with each of the fourth appearance features aw21, aw22, and aw23 to obtain the fifth merged feature corresponding to each fourth appearance feature. At this time, each fourth appearance feature corresponds to one fifth merged feature, and each third pose feature corresponds to three fifth merged features.
[0121] Similarly, taking the fourth pose feature az21 as an example, the fourth pose feature az21 is merged with each of the third appearance features aw11, aw12, and aw13 to obtain the sixth merged feature corresponding to each third appearance feature.
[0122] S1007: For each fifth merged feature corresponding to each third pose feature, the fifth merged feature is converted into a third synthetic image frame of the third sample pedestrian by the initial decoder; and for each sixth merged feature corresponding to each fourth pose feature, the sixth merged feature is converted into a fourth synthetic image frame of the third sample pedestrian by the initial decoder.
[0123] For example, taking the third pose feature az11 as an example, each fifth merged feature corresponding to the third pose feature az11 is input into the initial decoder. The initial decoder converts each fifth merged feature into a third composite image frame of the third template pedestrian. Each fifth merged feature corresponds to one third composite image frame. In this embodiment, each third pose feature corresponds to three fifth merged features, therefore each third pose feature corresponds to three third composite image frames.
[0124] The three third synthetic image frames corresponding to the third pose feature az11 contain: the third pose feature az11 and the fourth appearance feature aw21, the third pose feature az11 and the fourth appearance feature aw22, and the third pose feature az11 and the fourth appearance feature aw23, respectively.
[0125] The generation process parameters of the fourth composite image frame and the generation process of the third composite image frame are not described in detail in this application.
[0126] S1008: For each third synthetic image frame corresponding to each third pose feature, calculate the fourth pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the third synthetic image frame at each pixel point, to obtain the fourth pixel average of each third synthetic image frame; and for each fourth synthetic image frame corresponding to each fourth pose feature, calculate the fifth pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the fourth synthetic image frame at each pixel point, to obtain the fifth pixel average of each fourth synthetic image frame.
[0127] For example, taking the third pose feature az11 as an example, for the three third synthesized image frames corresponding to the third pose feature az11, the pixel difference between the third gait image frame A11 corresponding to the third pose feature az11 and the third synthesized image frame at each pixel point is calculated to obtain the pixel difference at each pixel point. Then, the average value of the pixel differences at all pixel points is calculated to obtain the fourth pixel average value. Each third synthesized image frame corresponds to a fourth pixel average value, and the three third synthesized image frames corresponding to the third pose feature az11 thus correspond to three fourth pixel average values.
[0128] The calculation process of the average value of the fifth pixel and the calculation process of the average value of the fourth pixel are not described in detail in this application.
[0129] S1009: Calculate the average of all fourth pixels corresponding to all third pose features and the average of all fifth pixels corresponding to all fourth pose features to obtain the average of the sixth pixel corresponding to the third sample pedestrian, and use the average of the sixth pixel as the second loss value; the second loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0130] In this embodiment, the first loss value and the second loss value can be used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder together until the first loss value and the second loss value converge, at which point training stops and the trained encoder and decoder are obtained.
[0131] In one possible implementation, after step S1001 is completed, the following steps S1011-S1018 can also be performed:
[0132] S1011: Input the seventh gait sequence of the fourth sample pedestrian into the initial encoder to obtain the fifth appearance feature and fifth pose feature corresponding to each seventh gait image frame in the seventh gait sequence.
[0133] For example, the seventh gait sequence B1 of the fourth sample pedestrian B includes three seventh gait image frames: B11, B12, and B13. The seventh gait sequence B1 is input into the initial encoder, which outputs the fifth appearance feature bw11 and the fifth posture feature bz11 corresponding to the seventh gait image frame B11, as well as the fifth appearance feature bw12 and the fifth posture feature bz12 corresponding to the seventh gait image frame B12, and the fifth appearance feature bw13 and the fifth posture feature bz13 corresponding to the seventh gait image frame B13.
[0134] S1012: For each third pose feature, merge the third pose feature with each fifth appearance feature to obtain multiple seventh merged features corresponding to the third pose feature; and for each fourth pose feature, merge the fourth pose feature with each fifth appearance feature to obtain multiple eighth merged features corresponding to the fourth pose feature.
[0135] For example, taking the third pose feature az11 as an example, the third pose feature az11 is merged with each of the fifth appearance features bw11, bw12 and bw13 respectively to obtain the three seventh merged features corresponding to the third pose feature az11.
[0136] Taking the fourth pose feature az21 as an example, the fourth pose feature az21 is merged with each of the fifth appearance features bw11, bw12 and bw13 respectively to obtain the three eighth merged features corresponding to the fourth pose feature az21.
[0137] S1013: For each seventh merged feature corresponding to each third pose feature, the seventh merged feature is converted into the fifth synthetic image frame of the third sample pedestrian by the initial decoder; and for each eighth merged feature corresponding to each fourth pose feature, the eighth merged feature is converted into the sixth synthetic image frame of the third sample pedestrian by the initial decoder.
[0138] For example, taking the third pose feature az11 as an example, for each seventh merged feature corresponding to the third pose feature az11, each seventh merged feature corresponding to the third pose feature az11 is input into the initial decoder, and the initial decoder converts the seventh merged feature into the fifth synthetic image frame of the third sample pedestrian. Here, each seventh merged feature corresponds to one fifth synthetic image frame, and each third pose feature corresponds to three fifth synthetic image frames.
[0139] The three fifth composite image frames corresponding to the third pose feature az11 respectively contain: the third pose feature az11 and the fifth appearance feature bw11, the third pose feature az11 and the fifth appearance feature bw12, and the third pose feature az11 and the fifth appearance feature bw13.
[0140] Taking the fourth pose feature az21 as an example, for each eighth merged feature corresponding to the fourth pose feature az21, each eighth merged feature corresponding to the fourth pose feature az21 is input into the initial decoder. The initial decoder converts the eighth merged feature into the sixth synthesized image frame of the third sample pedestrian. Each eighth merged feature corresponds to one sixth synthesized image frame, and each fourth pose feature corresponds to three sixth synthesized image frames.
[0141] The three sixth composite image frames corresponding to the fourth pose feature az21 respectively contain: the fourth pose feature az21 and the fifth appearance feature bw11, the fourth pose feature az21 and the fifth appearance feature bw12, and the fourth pose feature az21 and the fifth appearance feature bw13.
[0142] S1014: Input each fifth synthesized image frame into the initial encoder to obtain the sixth pose feature and sixth appearance feature corresponding to each fifth synthesized image frame; and input each sixth synthesized image frame into the initial encoder to obtain the seventh pose feature and seventh appearance feature corresponding to each sixth synthesized image frame.
[0143] For example, taking one of the fifth synthetic image frames corresponding to the third pose feature az11 as an example, the fifth synthetic image frame is input into the initial encoder to obtain the sixth pose feature az31 and the sixth appearance feature aw31 corresponding to the fifth synthetic image frame.
[0144] Taking one of the sixth synthetic image frames corresponding to the fourth pose feature az21 as an example, the sixth synthetic image frame is input into the initial encoder to obtain the seventh pose feature az41 and the seventh appearance feature aw41 corresponding to the sixth synthetic image frame.
[0145] S1015: For each sixth pose feature, merge the sixth pose feature with each third appearance feature to obtain multiple ninth merged features corresponding to the sixth pose feature; and for each seventh pose feature, merge the seventh pose feature with each third appearance feature to obtain multiple tenth merged features corresponding to the seventh pose feature.
[0146] S1016: For each ninth merged feature corresponding to each sixth pose feature, the ninth merged feature is converted into a seventh composite image frame of the third sample pedestrian by the initial decoder; and for each tenth merged feature corresponding to each seventh pose feature, the tenth merged feature is converted into an eighth composite image frame of the third sample pedestrian by the initial decoder.
[0147] S1017: For each seventh composite image frame corresponding to each sixth pose feature, calculate the seventh pixel average of the pixel difference between the third gait image frame corresponding to the sixth pose feature and the seventh composite image frame at each pixel point; and for each eighth composite image frame corresponding to each seventh pose feature, calculate the eighth pixel average of the pixel difference between the fourth gait image frame corresponding to the seventh pose feature and the eighth composite image frame at each pixel point.
[0148] For example, taking the sixth pose feature az31 as an example, for each seventh synthesized image frame corresponding to the sixth pose feature az31, the seventh pixel average value of the pixel difference between the third gait image frame A11 corresponding to the sixth pose feature az31 and the seventh synthesized image frame at each pixel point is calculated.
[0149] Taking the seventh pose feature az41 as an example, for each eighth synthesized image frame corresponding to the seventh pose feature az41, the average value of the eighth pixel difference between the fourth gait image frame A21 corresponding to the seventh pose feature az41 and the eighth synthesized image frame at each pixel point is calculated.
[0150] S1018: Calculate the average of all seventh pixel values corresponding to all sixth pose features and the average of all eighth pixel values corresponding to all seventh pose features to obtain the ninth pixel average value, which is used as the third loss value; the third loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0151] In this embodiment, the learnable parameters in the initial encoder and the learnable parameters in the initial decoder can be trained using a first loss value, a second loss value, and a third loss value. Training stops when the first loss value, the second loss value, and the trained encoder and decoder are obtained.
[0152] In one possible implementation, after performing step S1001, the following may also be added:
[0153] Calculate the average features of each third pose feature of the third sample pedestrian to obtain the first pose average feature; and calculate the average features of each fourth pose feature of the third sample pedestrian to obtain the second pose average feature.
[0154] The difference between the average features of the first pose and the average features of the second pose is calculated to obtain the fourth loss value; the fourth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0155] In this embodiment, the first loss value, the second loss value, the third loss value and the fourth loss value can be used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder together until the first loss value, the second loss value, the third loss value and the fourth loss value converge, and the training stops, thus obtaining the encoder and decoder after training.
[0156] In one possible implementation, after performing step S1011, the following may also be done:
[0157] For each third pose feature of the third sample pedestrian, the third pose feature is merged with each fifth appearance feature to obtain multiple eleventh merged features corresponding to the third pose feature.
[0158] The first InfoNCE loss value is calculated using the InfoNCE Loss function with the third pose feature, the third appearance feature, and the eleventh merged feature; and the second InfoNCE loss value is calculated using the InfoNCE Loss function with the fifth appearance feature, the fifth pose feature, and the eleventh merged feature.
[0159] The sum of the first InfoNCE loss value and the second InfoNCE loss value is determined as the fifth loss value; the fifth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0160] In this embodiment, the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value can be used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder together until the first loss value, the second loss value, the third loss value, the fourth loss value and the fifth loss value converge, and the training stops, thus obtaining the encoder and decoder after training.
[0161] Example 2:
[0162] Based on the same technical concept, this application also provides a device for generating gait recognition training data. Figure 3This illustration shows a schematic diagram of a gait recognition training data generation device provided in an embodiment of this application. Figure 3 As shown, the device includes:
[0163] The first input module 301 is used to input the first step state sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder to obtain the first pose feature and first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and to obtain the second pose feature and second appearance feature of the second sample pedestrian corresponding to each second gait image frame in the second gait sequence.
[0164] The first merging module 302 is configured to merge each first appearance feature with each second posture feature to obtain a plurality of first merged features corresponding to the first appearance feature; and to merge each second appearance feature with each first posture feature to obtain a plurality of second merged features corresponding to the second appearance feature.
[0165] The generation module 303 is used to input multiple first merged features corresponding to the same first appearance feature into a pre-trained decoder to generate a third gait sequence of the second sample pedestrian with the first appearance feature; and to input multiple second merged features corresponding to the same second appearance feature into the decoder to generate a fourth gait sequence of the first sample pedestrian with the second appearance feature.
[0166] Optionally, it also includes:
[0167] The acquisition module is used to acquire the first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian from the original training set before the first input module 301 inputs the first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian into the pre-trained encoder;
[0168] The addition module is used to add the third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian to the original training set after the generation module 303 generates the third gait sequence and the fourth gait sequence, to obtain a new training set; the new training set contains multiple gait sequences corresponding to each sample pedestrian, and each gait sequence corresponding to the same sample pedestrian contains its own appearance.
[0169] Optionally, it also includes:
[0170] The second input module is used to input the fifth and sixth gait sequences of the third sample pedestrian under different appearances into the initial encoder to be trained before the first input module 301 inputs the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, so as to obtain the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence.
[0171] The second merging module is used to merge each of the third posture features with each of the third appearance features corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and to merge each of the fourth posture features with each of the fourth appearance features corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature.
[0172] The first conversion module is used to convert each of the third merged features corresponding to each of the third pose features into a first synthetic image frame of the third sample pedestrian through an initial decoder to be trained, thereby obtaining a first synthetic image frame corresponding to each of the third merged features; and to convert each of the fourth merged features corresponding to each of the fourth pose features into a second synthetic image frame of the third sample pedestrian through the initial decoder.
[0173] The first calculation module is configured to, for each of the first synthetic image frames corresponding to the third pose feature, calculate the first average pixel value of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthetic image frame at each pixel point, to obtain the first average pixel value corresponding to each of the first synthetic image frames; and for each of the second synthetic image frames corresponding to the fourth pose feature, calculate the second average pixel value of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthetic image frame at each pixel point, to obtain the second average pixel value corresponding to each of the second synthetic image frames.
[0174] The second calculation module is used to calculate the average of all the first pixel averages corresponding to all the third pose features and the average of all the second pixel averages corresponding to all the fourth pose features, to obtain the third pixel average corresponding to the third sample pedestrian, and to use the third pixel average as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0175] Optionally, it also includes:
[0176] The third merging module is used to, after obtaining the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence from the second input module, merge the third pose feature with each of the fourth appearance features to obtain multiple fifth merged features corresponding to the third pose feature; and merge the fourth pose feature with each of the third appearance features to obtain multiple sixth merged features corresponding to the fourth pose feature.
[0177] The second conversion module is used to convert each fifth merged feature corresponding to each third pose feature into a third synthetic image frame of the third sample pedestrian through the initial decoder; and to convert each sixth merged feature corresponding to each fourth pose feature into a fourth synthetic image frame of the third sample pedestrian through the initial decoder.
[0178] The third calculation module is used to calculate, for each of the third synthetic image frames corresponding to the third pose feature, the fourth pixel average value of the pixel difference between the third gait image frame corresponding to the third pose feature and the third synthetic image frame at each pixel point, to obtain the fourth pixel average value corresponding to each of the third synthetic image frames; and to calculate, for each of the fourth synthetic image frames corresponding to the fourth pose feature, the fifth pixel average value of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the fourth synthetic image frame at each pixel point, to obtain the fifth pixel average value corresponding to each of the fourth synthetic image frames.
[0179] The fourth calculation module is used to calculate the average of all the fourth pixel averages corresponding to all the third pose features and the average of all the fifth pixel averages corresponding to all the fourth pose features, to obtain the sixth pixel average corresponding to the third sample pedestrian, and to use the sixth pixel average as the second loss value; the second loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0180] Optionally, it also includes:
[0181] The third input module is used to input the seventh gait sequence of the fourth sample pedestrian into the initial encoder to obtain the fifth appearance feature and fifth pose feature corresponding to each seventh gait image frame in the seventh gait sequence;
[0182] The fourth merging module is used to merge each of the third posture features with each of the fifth appearance features to obtain a plurality of seventh merged features corresponding to the third posture feature; and to merge each of the fourth posture features with each of the fifth appearance features to obtain a plurality of eighth merged features corresponding to the fourth posture feature.
[0183] The third conversion module is used to convert each of the seventh merged features corresponding to each of the third pose features into a fifth synthetic image frame of the third sample pedestrian through the initial decoder; and to convert each of the eighth merged features corresponding to each of the fourth pose features into a sixth synthetic image frame of the third sample pedestrian through the initial decoder.
[0184] The fourth input module is used to input each of the fifth synthesized image frames into the initial encoder to obtain the sixth pose feature and the sixth appearance feature corresponding to each of the fifth synthesized image frames; and to input each of the sixth synthesized image frames into the initial encoder to obtain the seventh pose feature and the seventh appearance feature corresponding to each of the sixth synthesized image frames.
[0185] The fifth merging module is used to merge each of the sixth posture features with each of the third appearance features to obtain multiple ninth merged features corresponding to the sixth posture feature; and to merge each of the seventh posture features with each of the third appearance features to obtain multiple tenth merged features corresponding to the seventh posture feature.
[0186] The fourth conversion module is used to convert each ninth merged feature corresponding to each sixth pose feature into a seventh composite image frame of the third sample pedestrian through the initial decoder; and to convert each tenth merged feature corresponding to each seventh pose feature into an eighth composite image frame of the third sample pedestrian through the initial decoder.
[0187] The fifth calculation module is used to calculate, for each of the seventh composite image frames corresponding to the sixth pose feature, the seventh pixel average of the pixel difference between the third gait image frame corresponding to the sixth pose feature and the seventh composite image frame at each pixel point; and to calculate, for each of the eighth composite image frames corresponding to the seventh pose feature, the eighth pixel average of the pixel difference between the fourth gait image frame corresponding to the seventh pose feature and the eighth composite image frame at each pixel point.
[0188] The sixth calculation module is used to calculate the average of the average values of all seventh pixels corresponding to all sixth pose features and the average values of all eighth pixels corresponding to all seventh pose features to obtain the average value of the ninth pixel, and use the average value of the ninth pixel as the third loss value; the third loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0189] Optionally, it also includes:
[0190] The seventh calculation module is used to calculate the average feature of each of the third posture features of the third sample pedestrian to obtain the first posture average feature; and to calculate the average feature of each of the fourth posture features of the third sample pedestrian to obtain the second posture average feature.
[0191] The eighth calculation module is used to calculate the difference between the average features of the first pose and the average features of the second pose to obtain a fourth loss value; the fourth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0192] Optionally, it also includes:
[0193] The sixth merging module is used to merge the third pose feature with each of the fifth appearance features for each of the third pose features of the third sample pedestrian to obtain multiple eleventh merged features corresponding to the third pose feature.
[0194] The ninth calculation module is used to calculate a first InfoNCE loss value using the third pose feature, the third appearance feature, and the eleventh merged feature through the InfoNCE Loss function; and to calculate a second InfoNCE loss value using the fifth appearance feature, the fifth pose feature, and the eleventh merged feature through the InfoNCE Loss function.
[0195] A determination module is used to determine the sum of the first InfoNCE loss value and the second InfoNCE loss value as a fifth loss value; the fifth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
[0196] Example 3:
[0197] Figure 4A schematic diagram of an electronic device provided in this application embodiment includes: a processor 401, a memory 402, and a bus 403. The memory 402 stores machine-readable instructions executable by the processor 401. When the electronic device runs the above-described information processing method, the processor 401 and the memory 402 communicate through the bus 403. The processor 401 executes the machine-readable instructions to perform the steps of the method described in Embodiment 1.
[0198] Example 4:
[0199] Embodiment 4 of this application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps described in Embodiment 1.
[0200] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, electronic devices, and computer-readable storage media described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0201] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0202] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0203] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0204] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0205] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The scope of protection of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of protection of the claims.
Claims
1. A method for generating gait recognition training data, characterized in that, include: The first step state sequence of the first sample pedestrian and the second step state sequence of the second sample pedestrian are input into the pre-trained encoder to obtain the first pose feature and the first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and the second pose feature and the second appearance feature of the second sample pedestrian corresponding to each second step state image frame in the second step state sequence. For each of the first appearance features, the first appearance feature is merged with each of the second pose features to obtain multiple first merged features corresponding to the first appearance feature; And for each second appearance feature, the second appearance feature is merged with each of the first pose features to obtain multiple second merged features corresponding to the second appearance feature; Multiple first merged features corresponding to the same first appearance feature are input into a pre-trained decoder to generate a third gait sequence of the second sample pedestrian with the same first appearance feature; and multiple second merged features corresponding to the same second appearance feature are input into the decoder to generate a fourth gait sequence of the first sample pedestrian with the same second appearance feature. Before inputting the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, the method further includes: The fifth and sixth gait sequences of the third sample pedestrian under different appearances are input into the initial encoder to be trained to obtain the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence. For each of the third posture features, the third posture feature is merged with each of the third appearance features corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the fourth appearance features corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature. For each of the third pose features corresponding to each of the third merged features, the third merged feature is converted into a first synthetic image frame of the third sample pedestrian by the initial decoder to be trained, so as to obtain the first synthetic image frame corresponding to each of the third merged features; and for each of the fourth pose features corresponding to each of the fourth merged features, the fourth merged feature is converted into a second synthetic image frame of the third sample pedestrian by the initial decoder. For each first synthesized image frame corresponding to each of the third pose features, calculate the first pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthesized image frame at each pixel point to obtain the first pixel average of each first synthesized image frame; and for each second synthesized image frame corresponding to each of the fourth pose features, calculate the second pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthesized image frame at each pixel point to obtain the second pixel average of each second synthesized image frame. The average of all first pixel values corresponding to all third pose features and the average of all second pixel values corresponding to all fourth pose features are calculated to obtain the third pixel average value corresponding to the third sample pedestrian, and the third pixel average value is used as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
2. The method according to claim 1, characterized in that, Before inputting the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, the method further includes: Obtain the first gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian from the original training set; After generating the third gait sequence and the fourth gait sequence, the method further includes: The third gait sequence of the second sample pedestrian and the fourth gait sequence of the first sample pedestrian are added to the original training set to obtain a new training set; the new training set contains multiple gait sequences corresponding to each sample pedestrian, and each gait sequence corresponding to the same sample pedestrian contains its own appearance.
3. The method according to claim 1, characterized in that, After obtaining the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence, the method further includes: For each of the third posture features, the third posture feature is merged with each of the fourth appearance features to obtain a plurality of fifth merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the third appearance features to obtain a plurality of sixth merged features corresponding to the fourth posture feature. For each fifth merged feature corresponding to each third pose feature, the fifth merged feature is converted into a third synthetic image frame of the third sample pedestrian by the initial decoder; and for each sixth merged feature corresponding to each fourth pose feature, the sixth merged feature is converted into a fourth synthetic image frame of the third sample pedestrian by the initial decoder. For each of the third synthetic image frames corresponding to each of the third pose features, calculate the fourth pixel average of the pixel difference between the third gait image frame corresponding to the third pose feature and the third synthetic image frame at each pixel point, to obtain the fourth pixel average of each of the third synthetic image frames; and for each of the fourth synthetic image frames corresponding to each of the fourth pose features, calculate the fifth pixel average of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the fourth synthetic image frame at each pixel point, to obtain the fifth pixel average of each of the fourth synthetic image frames; The average of the average of all fourth pixels corresponding to all third pose features and the average of all fifth pixels corresponding to all fourth pose features are calculated to obtain the average of the sixth pixel corresponding to the third sample pedestrian, and the average of the sixth pixel is used as the second loss value; the second loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
4. The method according to claim 3, characterized in that, The method further includes: The seventh gait sequence of the fourth sample pedestrian is input into the initial encoder to obtain the fifth appearance feature and fifth pose feature corresponding to each seventh gait image frame in the seventh gait sequence; For each of the third posture features, the third posture feature is merged with each of the fifth appearance features to obtain a plurality of seventh merged features corresponding to the third posture feature; and for each of the fourth posture features, the fourth posture feature is merged with each of the fifth appearance features to obtain a plurality of eighth merged features corresponding to the fourth posture feature. For each of the seventh merged features corresponding to each of the third pose features, the seventh merged feature is converted into a fifth composite image frame of the third sample pedestrian by the initial decoder; and for each of the eighth merged features corresponding to each of the fourth pose features, the eighth merged feature is converted into a sixth composite image frame of the third sample pedestrian by the initial decoder. Each of the fifth synthesized image frames is input into the initial encoder to obtain the sixth pose feature and the sixth appearance feature corresponding to each of the fifth synthesized image frames; and each of the sixth synthesized image frames is input into the initial encoder to obtain the seventh pose feature and the seventh appearance feature corresponding to each of the sixth synthesized image frames. For each of the sixth posture features, the sixth posture feature is merged with each of the third appearance features to obtain multiple ninth merged features corresponding to the sixth posture feature; and for each of the seventh posture features, the seventh posture feature is merged with each of the third appearance features to obtain multiple tenth merged features corresponding to the seventh posture feature. For each of the ninth merged features corresponding to the sixth pose feature, the initial decoder converts the ninth merged feature into the seventh composite image frame of the third sample pedestrian; and for each of the tenth merged features corresponding to the seventh pose feature, the initial decoder converts the tenth merged feature into the eighth composite image frame of the third sample pedestrian. For each of the seventh synthesized image frames corresponding to each of the sixth pose features, calculate the seventh pixel average of the pixel difference between the third gait image frame corresponding to the sixth pose feature and the seventh synthesized image frame at each pixel point; and for each of the eighth synthesized image frames corresponding to each of the seventh pose features, calculate the eighth pixel average of the pixel difference between the fourth gait image frame corresponding to the seventh pose feature and the eighth synthesized image frame at each pixel point. The average of all seventh pixel values corresponding to all sixth pose features and the average of all eighth pixel values corresponding to all seventh pose features are calculated to obtain the ninth pixel average value, which is used as the third loss value. The third loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
5. The method according to claim 3, characterized in that, The method further includes: Calculate the average feature of each of the third pose features of the third sample pedestrian to obtain the first pose average feature; and calculate the average feature of each of the fourth pose features of the third sample pedestrian to obtain the second pose average feature; The difference between the average features of the first pose and the average features of the second pose is calculated to obtain a fourth loss value; the fourth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
6. The method according to claim 4, characterized in that, The method further includes: For each of the third pose features of the third sample pedestrian, the third pose feature is merged with each of the fifth appearance features to obtain multiple eleventh merged features corresponding to the third pose feature; Using the InfoNCE Loss function, a first InfoNCE loss value is calculated using the third pose feature, the third appearance feature, and the eleventh merged feature; and using the InfoNCE Loss function, a second InfoNCE loss value is calculated using the fifth appearance feature, the fifth pose feature, and the eleventh merged feature. The sum of the first InfoNCE loss value and the second InfoNCE loss value is determined as the fifth loss value; the fifth loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
7. A device for generating gait recognition training data, characterized in that, include: The first input module is used to input the first step state sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder to obtain the first pose feature and the first appearance feature of the first sample pedestrian corresponding to each first step state image frame in the first step state sequence, and to obtain the second pose feature and the second appearance feature of the second sample pedestrian corresponding to each second gait image frame in the second gait sequence. The first merging module is used to merge each first appearance feature with each of the second pose features to obtain multiple first merged features corresponding to the first appearance feature. And for each second appearance feature, the second appearance feature is merged with each of the first pose features to obtain multiple second merged features corresponding to the second appearance feature; The generation module is used to input multiple first merged features corresponding to the same first appearance feature into a pre-trained decoder to generate a third gait sequence of the second sample pedestrian with the same first appearance feature; and to input multiple second merged features corresponding to the same second appearance feature into the decoder to generate a fourth gait sequence of the first sample pedestrian with the same second appearance feature. Also includes: The second input module is used to input the fifth and sixth gait sequences of the third sample pedestrian under different appearances into the initial encoder to be trained before the first input module inputs the first step gait sequence of the first sample pedestrian and the second gait sequence of the second sample pedestrian into the pre-trained encoder, so as to obtain the third appearance feature and third pose feature corresponding to each third gait image frame in the fifth gait sequence, and the fourth appearance feature and fourth pose feature corresponding to each fourth gait image frame in the sixth gait sequence. The second merging module is used to merge each of the third posture features with each of the third appearance features corresponding to other third posture features, to obtain multiple third merged features corresponding to the third posture feature; and to merge each of the fourth posture features with each of the fourth appearance features corresponding to other fourth posture features, to obtain multiple fourth merged features corresponding to the fourth posture feature. The first conversion module is used to convert each of the third merged features corresponding to each of the third pose features into a first synthetic image frame of the third sample pedestrian through an initial decoder to be trained, thereby obtaining a first synthetic image frame corresponding to each of the third merged features; and to convert each of the fourth merged features corresponding to each of the fourth pose features into a second synthetic image frame of the third sample pedestrian through the initial decoder. The first calculation module is configured to, for each of the first synthetic image frames corresponding to the third pose feature, calculate the first average pixel value of the pixel difference between the third gait image frame corresponding to the third pose feature and the first synthetic image frame at each pixel point, to obtain the first average pixel value corresponding to each of the first synthetic image frames; and for each of the second synthetic image frames corresponding to the fourth pose feature, calculate the second average pixel value of the pixel difference between the fourth gait image frame corresponding to the fourth pose feature and the second synthetic image frame at each pixel point, to obtain the second average pixel value corresponding to each of the second synthetic image frames. The second calculation module is used to calculate the average of all the first pixel averages corresponding to all the third pose features and the average of all the second pixel averages corresponding to all the fourth pose features, to obtain the third pixel average corresponding to the third sample pedestrian, and to use the third pixel average as the first loss value; the first loss value is used to train the learnable parameters in the initial encoder and the learnable parameters in the initial decoder.
8. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Model training method, clothing identification processing method, related device and terminal
CN114332940A
Model training method for gait recognition and gait recognition method and device
CN115205971A