A method and apparatus for extracting pedestrian features

CN118799917BActive Publication Date: 2026-09-04HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410928443.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2026-09-04
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是提供一种提取行人特征的方法和装置,其目的在于,提供了一种大规模自监督行人特征学习方法,结合原始行人图像和掩码行人图像,实现自监督行人重识别,解决了现有技术中由于有标注的训练数据不足,导致难以有效提取行人图像的局部细节特征、精度受限的问题

Benefits of technology

[0047]This invention performs two feature extractions on the same unlabeled pedestrian image using an initial pre-trained model. First, it converts the image into a sequence of original image blocks and extracts pedestrian features to obtain the teacher target vector. Second, it randomly masks the original words in the original image block sequence to obtain a masked image block sequence and extracts pedestrian features to obtain the student target vector. The teacher and student target vectors form a comparative relationship. Since the features learned by the student network are obtained by randomly masking the original words, the pre-trained model can express features at the smallest granularity. During the iterative training process combining the student and teacher networks, the pre-trained model learns to capture high-quality global pedestrian features and fine-grained local pedestrian features, thus avoiding reliance on labeled training data and effectively extracting local detail features of pedestrian images, achieving high-precision pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799917B_ABST
    Figure CN118799917B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of pedestrian re-identification, and provides a method and device for extracting pedestrian features. An unlabeled pedestrian image is input into an initial pre-training model, and is converted into an original picture block sequence; a teacher network of the initial pre-training model is used to extract pedestrian features in the original picture block sequence, and a teacher target vector is obtained; original word units in the original picture block sequence are randomly masked to obtain a mask picture block sequence; a student network of the initial pre-training model is used to extract pedestrian features in the mask picture block sequence, and a student target vector is obtained; network parameters of the initial pre-training model are optimized according to the teacher target vector and the student target vector, and a final target pre-training model is obtained through iterative training, so that the target pre-training model is used to extract pedestrian features. The application solves the problem that, in the prior art, due to insufficient labeled training data, it is difficult to effectively extract local detailed features of a pedestrian image, and the precision is limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pedestrian re-identification technology, and in particular to a method and apparatus for extracting pedestrian features. Background Technology

[0002] Person re-identification (Person ReID) enables effective cross-camera retrieval of pedestrians in contactless and non-cooperative scenarios, and is widely used in public safety, video surveillance, and other fields, demonstrating significant application value. Person re-identification requires learning visual features from pedestrian images that can distinguish the identities of different pedestrians, and overcoming challenges such as severe occlusion, appearance changes, shape changes, and viewpoint changes. It is an important and highly challenging computer vision problem.

[0003] Currently, solutions for pedestrian re-identification based on publicly available datasets and using supervised learning methods are relatively mature, with most involving metric learning based on features extracted by backbone networks. However, pedestrian re-identification faces the challenge of insufficient labeled training data: due to the high cost and difficulty of collecting and labeling pedestrian images, existing publicly available datasets are generally small, making it difficult for backbone network-based metric learning solutions to effectively extract local detail features from pedestrian images, thus limiting the accuracy of pedestrian re-identification.

[0004] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method and apparatus for extracting pedestrian features. The purpose is to provide a large-scale self-supervised pedestrian feature learning method that combines the original pedestrian image and the masked pedestrian image to achieve self-supervised pedestrian re-identification. This solves the problem in the prior art that the lack of labeled training data makes it difficult to effectively extract local detail features of pedestrian images and limits accuracy.

[0006] The present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides a method for extracting pedestrian features, comprising:

[0008] Unlabeled pedestrian images are input into an initial pre-trained model to convert them into a sequence of raw image patches. The teacher network of the initial pre-trained model is used to extract pedestrian features from the sequence of raw image patches to obtain the teacher target vector.

[0009] The original words in the original image block sequence are randomly masked to obtain a masked image block sequence; the student network of the initial pre-trained model is used to extract pedestrian features from the masked image block sequence to obtain the student target vector;

[0010] Optimize the network parameters of the initial pre-trained model based on the teacher's target vector and the student's target vector;

[0011] The above process is iterated until the change in the loss function of the initial pre-trained model is less than a preset convergence threshold, thus obtaining the target pre-trained model, which is then used to extract pedestrian features.

[0012] Furthermore, the original word element includes a classification original word element and at least one input original word element;

[0013] The step of randomly masking the original words in the original image block sequence to obtain the masked image block sequence includes:

[0014] The original lexical units of the classification are used as the classification mask lexical units;

[0015] From the at least one input original word, a preset number of input original words are selected, and the selected input original words are masked; based on the masked input original words and the unmasked input original words, the input image block sequence corresponding to the at least one input original word is obtained;

[0016] Based on the classification mask terms and the input image block sequence, a mask image block sequence is obtained.

[0017] Further, the step of selecting a preset number of input original words from the at least one input original word word, and masking the selected input original word words; obtaining the input image block sequence corresponding to the at least one input original word word based on the masked input original word words and the unmasked input original word words includes:

[0018] For the at least one input original word, the corresponding mask amount is set to a first preset value to select that the input original word be masked; the corresponding mask amount is set to a second preset value to select that the input original word be not masked.

[0019] When the mask amount of the input original word is a first preset value, the input original word is replaced with a mask to obtain the masked input original word; when the mask amount of the input original word is a second preset value, the input original word is treated as the unmasked input original word.

[0020] Furthermore, the expression for the masked image block sequence is:

[0021] in, For classification mask words, Given a sequence of image patches as input, The i-th input primitive word in the input image block sequence; x [mask] For masking, m i This is the mask value.

[0022] Furthermore, the student target vector includes a representative student target vector and at least one ordinary student target vector;

[0023] The student network using the initial pre-trained model extracts pedestrian features from the masked image block sequence to obtain the student target vector, which includes:

[0024] The student network is used to extract local pedestrian features from the classification mask words in the masked image block sequence to obtain the first representative feature; the first representative feature is mapped to the target vector space to obtain the representative student target vector;

[0025] The student network is used to predict the local pedestrian features of the masked original input words in the masked image block sequence, and the local pedestrian features of the unmasked original input words are extracted to obtain the second representative features; the second representative features are mapped to the target vector space to obtain the ordinary student target vector.

[0026] Furthermore, the teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector; the student target vector includes a representative student target vector and at least one ordinary student target vector; the pre-training loss function includes a mask reconstruction loss function and a global loss function;

[0027] The step of optimizing the network parameters of the initial pre-trained model based on the teacher's target vector and the student's target vector includes:

[0028] The mask reconstruction loss function is used to compare the at least one ordinary teacher target vector with the at least one ordinary student target vector to obtain the mask reconstruction loss value;

[0029] The global loss function is used to compare the representative teacher target vector with the representative student target vector to obtain the global loss value;

[0030] The product of the mask reconstruction loss value and the corresponding weight value is determined as the first intermediate value; the product of the global loss value and the corresponding weight value is determined as the second intermediate value; the sum of the first intermediate value and the second intermediate value is determined as the pre-training loss value, so as to calculate the pre-training loss value of this iteration through the pre-training loss function;

[0031] Based on the pre-training loss value, the network parameters of the initial pre-trained model are updated, and the next iteration is performed according to the updated network parameters.

[0032] Furthermore, the expression for the pre-training loss function is: L = λ1L d +λ2L m ; among which, L d Let L be the global loss function, and λ1 be the corresponding weight value; m Let λ2 be the loss function for mask reconstruction, and λ2 be the corresponding weight value.

[0033] The expression for the global loss function is: Among them, y [cls] To represent the teacher's target vector, Let P(x) represent the student's target vector, and let P(x) be the preset loss function.

[0034] The expression for the mask reconstruction loss function is: Where, for n target vectors of ordinary teachers and n target vectors of ordinary students, y i [patchs] Let i be the target vector of the i-th ordinary teacher. Let m be the target vector of the i-th ordinary student. i This is the mask value.

[0035] Furthermore, the original word unit includes a classification original word unit and at least one input original word unit; the teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector;

[0036] The process involves inputting unlabeled pedestrian images into an initial pre-trained model to convert the unlabeled pedestrian images into a sequence of raw image patches; then, using the teacher network of the initial pre-trained model, pedestrian features are extracted from the raw image patch sequence to obtain the teacher target vector, which includes:

[0037] The unlabeled pedestrian image is segmented into multiple image blocks; each image block is mapped to a patch word; from the patch words corresponding to the multiple image blocks, one patch word is selected as the original classification word, and the remaining patch words are used as the input original words to obtain the original image block sequence.

[0038] The teacher network is used to extract global pedestrian features from the original word units classified in the original image block sequence to obtain the third representative feature; the third representative feature is mapped to the target vector space to obtain the representative teacher target vector.

[0039] The teacher network is used to extract global pedestrian features from the input original words in the original image block sequence to obtain the fourth representative feature; the fourth representative feature is mapped to the target vector space to obtain the ordinary teacher target vector.

[0040] Secondly, the present invention also provides an apparatus for extracting pedestrian features, used to implement the method for extracting pedestrian features described in the first aspect, the apparatus comprising:

[0041] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor for performing the method for extracting pedestrian features as described in the first aspect.

[0042] Thirdly, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions that are executed by one or more processors to perform the method for extracting pedestrian features as described in the first aspect.

[0043] Fourthly, a chip is provided, comprising: a processor and an interface for calling and running a computer program stored in a memory, performing a method for extracting pedestrian features as described in the first aspect.

[0044] Fifthly, a computer program product containing instructions is provided, which, when executed on a computer or processor, cause the computer or processor to perform a method for extracting pedestrian features as described in the first to fourth aspects and any one thereof.

[0045] In a sixth aspect, a method system for extracting pedestrian features is provided, comprising an apparatus for extracting pedestrian features as described in the second aspect, and using the method for extracting pedestrian features as described in the first aspect to complete the interaction of the apparatus for extracting pedestrian features in the second aspect.

[0046] Unlike existing technologies, the present invention has at least the following beneficial effects:

[0047] This invention performs two feature extractions on the same unlabeled pedestrian image using an initial pre-trained model. First, it converts the image into a sequence of original image blocks and extracts pedestrian features to obtain the teacher target vector. Second, it randomly masks the original words in the original image block sequence to obtain a masked image block sequence and extracts pedestrian features to obtain the student target vector. The teacher and student target vectors form a comparative relationship. Since the features learned by the student network are obtained by randomly masking the original words, the pre-trained model can express features at the smallest granularity. During the iterative training process combining the student and teacher networks, the pre-trained model learns to capture high-quality global pedestrian features and fine-grained local pedestrian features, thus avoiding reliance on labeled training data and effectively extracting local detail features of pedestrian images, achieving high-precision pedestrian re-identification. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0049] Figure 1 This is a flowchart illustrating a method for extracting pedestrian features according to an embodiment of the present invention;

[0050] Figure 2 This is a schematic diagram of the network structure of a pre-trained model according to an embodiment of the present invention.

[0051] Figure 3 This is a flowchart illustrating step 10 provided in an embodiment of the present invention;

[0052] Figure 4 This is a flowchart illustrating a pedestrian image provided in an embodiment of the present invention;

[0053] Figure 5 This is a schematic diagram of a primitive word element provided in an embodiment of the present invention;

[0054] Figure 6 This is a flowchart illustrating step 20 provided in an embodiment of the present invention;

[0055] Figure 7 This is a schematic diagram of a mask provided in an embodiment of the present invention;

[0056] Figure 8 This is a flowchart illustrating step 202a provided in an embodiment of the present invention;

[0057] Figure 9This is a flowchart illustrating another step 20 provided in an embodiment of the present invention;

[0058] Figure 10 This is a flowchart illustrating step 30 provided in an embodiment of the present invention;

[0059] Figure 11 This is a schematic diagram illustrating a local feature clustering method provided in an embodiment of the present invention;

[0060] Figure 12 This is a schematic diagram illustrating a self-attention map display of a complex background provided by an embodiment of the present invention;

[0061] Figure 13 This is a schematic diagram illustrating an example of correlation analysis of different image features of the same pedestrian provided in an embodiment of the present invention;

[0062] Figure 14 This is a schematic diagram of the architecture of a device for extracting pedestrian features provided in an embodiment of the present invention. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0064] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0065] Unless the context otherwise requires, throughout the specification and claims, the term "comprising" is interpreted as openly inclusive, meaning "including, but not limited to." In the description of the specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this disclosure. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics mentioned may be included in any suitable manner in any one or more embodiments or examples; that is, although they may be incorporated into embodiments or examples using the above terms for reasons such as order and position, it does not limit them to be incorporated in combination by a single embodiment or example.

[0066] In the description of this invention, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this disclosure and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this disclosure.

[0067] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of embodiments of this disclosure, unless otherwise stated, "a plurality of" means two or more. Furthermore, for example, the description may use the prefix "A" or "B" to describe the same type of nouns as two independent entities. In this case, the corresponding features defined with "A" and "B" are used only to distinguish between similar entities and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features.

[0068] In describing some embodiments, the terms "coupled," "coupled," and "connected," and their derivative expressions, may be used. For example, the term "connected" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact with each other. Similarly, the term "coupled" may be used in describing some embodiments to indicate that two or more components have direct physical or electrical contact. However, the terms "connected" or "coupled" may also refer to two or more components that do not have direct contact with each other but still cooperate or interact with each other, such as "optical coupling," "wireless connection," etc. The embodiments disclosed herein are not necessarily limited to the scope of this invention.

[0069] In the description of this invention, the expression “A and / or B” (where A and B are used to formally represent specific features) will be used. The corresponding expression includes the following three combinations: only A, only B, and a combination of A and B.

[0070] As used in this invention, “about,” “approximately,” or “approximately” includes the stated value and the average value within an acceptable range of deviation from a particular value, wherein the acceptable range of deviation is determined by a person skilled in the art taking into account the measurement under discussion and the error associated with the measurement of the particular quantity (i.e., the limitations of the measurement system).

[0071] Example 1:

[0072] Traditional person re-identification schemes mostly rely on metric learning based on features extracted from a backbone network (e.g., residual networks). Due to the high cost and difficulty of collecting and labeling person re-identification datasets, and the generally small size of existing public datasets, the pre-training of the backbone network is crucial. Before the advent of self-supervised learning algorithms such as contrastive learning, backbone networks were primarily pre-trained for classification on the ImageNet dataset, followed by supervised fine-tuning on person re-identification datasets to achieve high accuracy and generalization ability. However, there are significant differences between the ImageNet dataset and person re-identification datasets. The ImageNet dataset contains images of a thousand categories, but only one type of person re-identification data—person images. Therefore, models pre-trained on the ImageNet dataset can extract category difference features, but struggle to effectively extract fine-grained individual features, making it unsuitable for the person re-identification problem.

[0073] Since pedestrian re-identification is an open problem addressing the fine-grained differences between pedestrians, its supervised training samples are only a subset of the pedestrian classification dataset, and the test pedestrians generally do not appear in the training subset. Therefore, the algorithm's generalization ability is crucial to its accuracy. However, due to the difficulty in collecting and labeling pedestrian re-identification datasets, currently available supervised training sets are generally small, making it difficult to meet the requirements for algorithm generalization. Furthermore, because the differences between different pedestrians are subtle, compared to classification problems that perform well using the ImageNet dataset, the differences between different pedestrians are much smaller than the differences between different classifications. Therefore, pre-trained models based on ImageNet classification and pedestrian re-identification schemes using feature extraction backbone networks are gradually being replaced by pre-trained models based on self-supervised learning from large-scale pedestrian data.

[0074] In recent years, with the development of computer vision and self-supervised learning techniques, the performance of person re-identification based on self-supervised pre-training has been significantly improved. For the person re-identification problem, capturing the subtle differences between different pedestrians—that is, the model's ability to extract fine-grained local features—is key to improving re-identification accuracy. Models based on certain algorithms possess strong global context-related feature representation capabilities, which can improve person re-identification performance to some extent; however, these models are better at extracting context-related global features and struggle to focus on local human features. Person re-identification requires extracting highly discriminative fine-grained local features of pedestrians to achieve further breakthroughs in recognition accuracy.

[0075] Furthermore, addressing the significant difference between pre-training datasets and pedestrian re-identification datasets, existing technologies utilize the unlabeled pedestrian dataset LUPerson. On this dataset, it was first demonstrated that self-supervised pre-trained models based on Convolutional Neural Networks (CNNs) exhibited excellent performance in pedestrian re-identification. Subsequently, some models were tested on the LUPerson dataset using self-supervised learning algorithms, achieving good results at the time. However, these self-supervised learning algorithms rely on comparing the classification consistency between local and global images generated by different data augmentations of the same image. While they effectively capture inter-class differences in multi-class datasets like ImageNet, they lose fine-grained differences in pedestrian data for single-class pedestrian datasets. For example, local images of different pedestrians may be very similar and belong to the same category, but the self-supervised comparison of these models will classify them into different categories. Similarly, local images of the same pedestrian from different angles that belong to the same category will be separated by the comparison of these models.

[0076] Therefore, models based on new algorithms emerged, proposing a self-supervised pre-training algorithm designed for pedestrian re-identification tasks. Building upon the aforementioned scheme, pedestrian images were divided into several blocks according to human body structure from top to bottom. Self-learning was performed between blocks and between blocks and the global image, thereby enhancing the pre-trained model's ability to represent local pedestrian features and achieving the highest accuracy in the industry at the time. However, this type of model has two problems: 1) There is an alignment dependency problem, that is, in pedestrian images with complex backgrounds, irregular alignment, or even occlusion and incompleteness, this block division method is too mechanical and prone to misclassification; 2) Even with fewer blocks, it is still difficult to fully express local fine-grained features. For example, the optimal experiment of the most accurate model is to divide the pedestrian into 3 blocks, but the pedestrian is full of detailed clues from head to toe, and dividing it into only 3 blocks is far from meeting the needs of pedestrian re-identification to capture local fine-grained features.

[0077] To solve the above problems, such as Figure 1 As shown, this embodiment of the invention provides a method for extracting pedestrian features, including:

[0078] Step 10: Input the unlabeled pedestrian image into the initial pre-trained model to convert the unlabeled pedestrian image into a sequence of original image patches; use the teacher network of the initial pre-trained model to extract pedestrian features from the sequence of original image patches to obtain the teacher target vector.

[0079] This invention introduces a self-supervised learning (SSL) method based on masked image modeling (MIM) into person re-identification. By combining masked image modeling and discriminative contrastive learning (i.e., a self-supervised learning method), discriminative features are learned from a large-scale unlabeled person re-identification dataset through large-scale self-supervised pre-training, thereby effectively extracting high-quality global and local features. Then, the global and local features are combined for supervised fine-tuning training in the person re-identification task.

[0080] To facilitate comparative learning, embodiments of the present invention preprocess unlabeled pedestrian images, converting them into a series of processable vectors (i.e., sequences of original image patches), and inputting them into a teacher network, which extracts the global features contained in the unlabeled pedestrian images. In an optional embodiment, the teacher network can be implemented based on the Vision Transformer (ViT) algorithm; the ViT algorithm can utilize a self-attention mechanism to model global relationships in pedestrian images.

[0081] Step 20: Randomly mask the original words in the original image block sequence to obtain a masked image block sequence; use the student network of the initial pre-trained model to extract pedestrian features from the masked image block sequence to obtain the student target vector.

[0082] To facilitate comparative learning, this embodiment of the invention randomly masks the original words in the original image block sequence to obtain a masked image block sequence, which is then input into a student network. The student network learns and predicts the randomly masked local fine-grained features. In an optional embodiment, the student network can be implemented based on the ViT model.

[0083] Masking involves overlaying a mask onto the vectors in the original image patch sequence, thereby blocking specific elements in the masked image patch sequence. The masked locations contain no features at all, and the student network learns to predict the features at the masked locations. The original lexical units are vector representations of a certain image patch in the unlabeled pedestrian image, which is equivalent to randomly masking the unlabeled pedestrian image patch by patch.

[0084] Step 30: Optimize the network parameters of the initial pre-trained model based on the teacher target vector and the student target vector.

[0085] Step 40: Iterate the above process until the change value of the loss function of the initial pre-trained model is less than the preset convergence threshold to obtain the target pre-trained model, so as to extract pedestrian features using the target pre-trained model.

[0086] The loss change value is the amount of change in the loss value calculated according to the pre-trained loss function; the preset convergence threshold is selected by those skilled in the art based on the specific application scenario.

[0087] In each iteration of this invention, the learning results of the student network and the teacher network (i.e., the teacher target vector and the student target vector) are combined to calculate the pre-training loss function until the network parameters converge and the training is completed, resulting in a well-trained target pre-trained model. This model can then be used to extract pedestrian features and perform subsequent pedestrian re-identification work.

[0088] In one alternative embodiment, the network can be pre-trained on the LUPerson dataset with ViT-S and ViT-B as the backbone, and then supervised fine-tuning can be performed on four mainstream public datasets for person re-identification: MSMT17, Market1501, DukeMTMC-reID, and Occluded-Duke.

[0089] This invention performs two feature extractions on the same unlabeled pedestrian image using an initial pre-trained model. First, the image is converted into a sequence of original image blocks, and pedestrian features are extracted to obtain the teacher target vector. Second, the original words in the original image block sequence are randomly masked to obtain a masked image block sequence, and pedestrian features are extracted to obtain the student target vector. The teacher and student target vectors form a comparative relationship. Since the features learned by the student network are obtained by randomly masking the original words, the pre-trained model can express features at the smallest granularity. During the iterative training process combining the student and teacher networks, the pre-trained model learns to capture high-quality global pedestrian features and fine-grained local pedestrian features, thus avoiding reliance on labeled training data and effectively extracting local detail features of pedestrian images, achieving high-precision pedestrian re-identification.

[0090] To illustrate the process of extracting global features from the teacher network, the original lexical units include a classification original lexical unit and at least one input original lexical unit; the teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector.

[0091] The network structure of the pre-trained model in this embodiment of the invention is as follows: Figure 2 As shown; the original word units in the original image block sequence include one classification word unit and at least one input word unit. After being input into the teacher network, they will respectively produce corresponding results (i.e., pedestrian features); where z [cls] To classify the results corresponding to the original word units, z [patchs] For at least one input original word, there is at least one result.

[0092] like Figure 3 As shown, step 10 includes:

[0093] Step 101: Segment the unlabeled pedestrian image into multiple image blocks; map each image block to a patch word; select one patch word from the patch words corresponding to the multiple image blocks as the original classification word, and use the remaining patch words as the input original words to obtain the original image block sequence.

[0094] To process unlabeled pedestrian images into a form that facilitates computation by the pre-trained model, we first process the unlabeled pedestrian images I∈R... h×w×c Encode the image to obtain the corresponding feature vector, where h×w is the resolution of the pedestrian image and c is the number of channels of the pedestrian image, usually 3, representing the three channels of red, green and blue (RGB).

[0095] like Figure 4 As shown, a pedestrian image is a continuous whole. To facilitate global relationship modeling, it needs to be converted into tokens, similar to word segmentation in natural language processing, resulting in multiple image patch tokens. Each token represents a local region in the pedestrian image and can be seen as an abstract representation of that local region. This embodiment of the invention performs block projection tokenization on the pedestrian image, dividing it into fixed-size, easily processed image blocks. For example, dividing it into blocks of p×p pixels can result in n = hw / p blocks. 2 Each image patch.

[0096] Each image patch consists of a set of pixels. In this embodiment of the invention, the image patch is used as the basic unit for generating words, and is mapped to patch words through a linear transformation. For example... Figure 5 As shown, each image patch can be represented as a patch term. Furthermore, to facilitate the extraction of pedestrian features based on the pre-trained model of this invention and to implement various subsequent pedestrian re-identification schemes, one patch term is selected from the obtained patch terms as the original classification term. This original classification term is learnable. The position of the original classification term in the original image patch sequence is selected by those skilled in the art according to the specific application scenario. In optional embodiments, such as... Figure 5 As shown, the original classification term is located at the first position in the original image block sequence. When the subsequent person re-identification scheme is implemented through a classification task, the selected original classification term is equivalent to a classification token. The teacher target vector after the original classification term passes through the teacher network can be used as the basis for classification.

[0097] Finally, the tokenized result is represented as the original image patch sequence as follows:

[0098] X = (x [cls] ;x [patchs] )=(x [cls] ;x1;...;x n )

[0099] Where, x [cls] To classify the original word units, x [patchs] For the input raw word units, x [patchs] =x1;...;x n n

[0100] Input the number of original words in the original image block sequence.

[0101] Step 102: Use the teacher network to extract global pedestrian features of the original word units classified in the original image block sequence to obtain the third representative feature; map the third representative feature to the target vector space to obtain the representative teacher target vector.

[0102] To further extract features, after obtaining the output through the teacher network, it is necessary to map the corresponding output to the target vector space. In an optional embodiment, a multilayer perceptron (MLP) can be used to map the third representative features to the target vector space; an MLP is a basic artificial neural network model. The corresponding expression is as follows:

[0103] MLP(z [cls] )=y [cls]

[0104] Among them, MLP(z) [cls] ) indicates that the third representative feature is mapped to the target vector space using an MLP, z [cls] As the third representative feature, y [cls] Let represent the teacher's target vector.

[0105] Step 103: Use the teacher network to extract global pedestrian features from the input original words in the original image block sequence to obtain the fourth representative feature; map the fourth representative feature to the target vector space to obtain the ordinary teacher target vector.

[0106] Similarly, in an alternative embodiment, an MLP can be used to map the fourth representative feature to the target vector space.

[0107] The expression for the output of the original image patch sequence corresponding to the teacher network is:

[0108] Z = (z [cls] ;z [patchs])=(z [cls] ;z1;...;z n )

[0109] Among them, z [patchs] As the fourth representative feature, z [patchs] =z1;...;z n .

[0110] The expression for the teacher's target vector is as follows:

[0111] Y = (MLP1(z) [cls] MLP2(z) [patchs] ))=(y [cls] ;y [patchs] )

[0112] Among them, MLP1(z [cls] ) indicates that the third representative feature is mapped to the target vector space using the first MLP network, MLP2(z) [patchs] ) indicates that the third representative feature is mapped to the target vector space using a second MLP network, y [cls] =MLP1(z [cls] ), y [patchs] =MLP2(z [patchs] ).

[0113] The following explains the process of extracting pedestrian features from the student network:

[0114] To illustrate the process of random masking, such as Figure 6 As shown, the original word unit includes a classification original word unit and at least one input original word unit;

[0115] In step 20, the random masking of the original words in the original image block sequence to obtain the masked image block sequence includes:

[0116] Step 201a: Use the original classification lexical units as classification mask lexical units.

[0117] It should be noted that since the output of the classification masking terminology is used for subsequent related tasks, it is generally not masked.

[0118] Step 202a: Select a preset number of input original words from the at least one input original word, and mask the selected input original words; based on the masked input original words and the unmasked input original words, obtain the input image block sequence corresponding to the at least one input original word.

[0119] The preset number is selected by those skilled in the art based on the specific usage scenario; in an optional embodiment, the preset number can be 1, i.e., as shown below. Figure 7 As shown, for such Figure 4 In the pedestrian image shown, multiple image patches are selected, and only one is masked. Because this invention performs random masking on an image patch basis, the final target pre-trained model can represent features at the smallest granularity (i.e., image patch).

[0120] Step 203a: Obtain the mask image block sequence based on the classification mask words and the input image block sequence.

[0121] The selected original input tokens are not visible in the masked image block sequence.

[0122] To explain the process of achieving concealment in detail, such as Figure 8 As shown, step 202a includes:

[0123] Step 2021: For the at least one input original word, set the corresponding mask amount to a first preset value to select to mask the input original word; set the corresponding mask amount to a second preset value to select not to mask the input original word.

[0124] The first preset value and the second preset value are selected by those skilled in the art according to the specific use scenario. The first preset value and the second preset value are different. In an optional embodiment, the first preset value can be 1 and the second preset value can be 0.

[0125] The expression for the masked image block sequence is:

[0126] in, For classification mask words, Given a sequence of image patches as input, This is the i-th original word in the input image block sequence. The classification mask word is a learnable word (token).

[0127] Step 2022: When the mask amount of the input original word is a first preset value, the input original word is replaced with a mask to obtain the masked input original word; when the mask amount of the input original word is a second preset value, the input original word is used as the unmasked input original word.

[0128] In an optional embodiment, the above process is expressed as follows:

[0129] x [mask] For masking, m i This is the mask value.

[0130] To illustrate the process of obtaining the student's target vector, as follows: Figure 9As shown, the student target vector includes a representative student target vector and at least one ordinary student target vector; in step 20, the student network using the initial pre-trained model extracts pedestrian features from the masked image block sequence to obtain the student target vector, which includes:

[0131] Step 201b: Use the student network to extract local pedestrian features of classification mask words in the masked image block sequence to obtain the first representative feature; map the first representative feature to the target vector space to obtain the representative student target vector.

[0132] Similarly, for the teacher network, in order to further extract features, after obtaining the output through the student network, it is also necessary to map the corresponding output to the target vector space. In an optional embodiment, an MLP can be used to map the first representative features to the target vector space.

[0133] Step 202b: Use the student network to predict the local pedestrian features of the masked input original words in the masked image block sequence, extract the local pedestrian features of the unmasked input original words to obtain the second representative features; map the second representative features to the target vector space to obtain the ordinary student target vector.

[0134] Similarly, in an alternative embodiment, an MLP can be used to map the second representative features to the target vector space. The expression for the student target vector is as follows:

[0135]

[0136] in, As the first representative feature, As the second representative characteristic, This indicates that the first MLP network is used to map the first representative feature to the target vector space. This indicates that a second MLP network is used to map the second representative feature to the target vector space.

[0137] This invention combines contrastive learning and masked image modeling, which can effectively solve the problem of learning fine-grained visual features of pedestrians in the case of severe occlusion and misalignment of pedestrian images. The following describes the design of the corresponding pre-training loss function when combining the two in pedestrian re-identification scenarios.

[0138] The teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector; the student target vector includes a representative student target vector and at least one ordinary student target vector; the pre-training loss function includes a mask reconstruction loss function and a global loss function.

[0139] like Figure 10 As shown, step 30 includes:

[0140] Step 301: Use the mask reconstruction loss function to compare the at least one ordinary teacher target vector with the at least one ordinary student target vector to obtain the mask reconstruction loss value.

[0141] The expression for the mask reconstruction loss function is: Where, for n target vectors of ordinary teachers and n target vectors of ordinary students, y i [patchs] Let i be the target vector of the i-th ordinary teacher. Let m be the target vector of the i-th ordinary student. i This is the mask value. In an optional embodiment, P(x) = softmax(x).

[0142] Step 302: Use the global loss function to compare the representative teacher target vector with the representative student target vector to obtain the global loss value.

[0143] The expression for the global loss function is: Among them, y [cls] To represent the teacher's target vector, Let P(x) represent the student's target vector, and let P(x) be the preset loss function.

[0144] Step 303: Determine the product of the mask reconstruction loss value and the corresponding weight value as the first intermediate value; determine the product of the global loss value and the corresponding weight value as the second intermediate value; determine the sum of the first intermediate value and the second intermediate value as the pre-training loss value, so as to calculate the pre-training loss value of this iteration through the pre-training loss function.

[0145] The expression for the pre-training loss function is: L = λ1L d +λ2L m ; among which, L d Let L be the global loss function, and λ1 be the corresponding weight value; m Let λ be the loss function for mask reconstruction, and λ2 be the corresponding weight value.

[0146] Step 304: Based on the pre-training loss value, update the network parameters of the initial pre-trained model, and perform the next iteration according to the updated network parameters.

[0147] It is worth noting that this invention introduces mask image modeling technology into the self-supervised feature learning of pedestrian re-identification, and combines it with contrastive learning to solve the problem of limited accuracy caused by the small number of pedestrian image blocks in the prior art. It has the ability to extract fine-grained features of key human body parts of pedestrians and surrounding objects.

[0148] The pre-trained model trained using the pedestrian feature extraction method of this invention has the characteristics of self-supervision, scalability, and strong generalization ability, and overcomes the problem of difficult labeling in supervised pedestrian re-identification. It has achieved state-of-the-art results on mainstream public datasets such as MSMT (Multi-Scene Multi-Time)17, Market1501, Duke MTMC-reID (Multi-Tracking Multi-Camera ReIDentification), and Occluded-Duke. For example, the model pre-trained with the ViT-B / 16 backbone network has a mean average precision (mAP) of up to 80.8 on MSMT17, which is 9.0 higher than the latest algorithm in the prior art; the rank-1 accuracy is up to 92.0, which is 3.8 higher than the latest algorithm in the prior art. The rank-1 accuracy measures the proportion of the highest probability (or most similar) category in the model's prediction results that is the same as the true label.

[0149] Furthermore, experiments were conducted on subsequent related tasks in this embodiment of the invention, using pedestrian features extracted by the target pre-trained model. To demonstrate the effectiveness of the target pre-trained model in this embodiment of the invention, the visualization results of subsequent related tasks and experiments are described below:

[0150] The experiment was based on the LUPerson dataset, using a pre-trained ViT-S / 16 model as the teacher network and student network, and performed visualization analysis on the MSMT17 dataset, including feature clustering analysis, self-attention view analysis, and feature correlation analysis.

[0151] like Figure 11 The figure shows the target vector y for the teacher. [patchs] Example image showing the results of local feature clustering analysis; Figure 11 Each subgraph in the diagram represents a cluster ( Figure 11 The CCP contains 6 sub-images, and the red dots represent the positions of the image blocks to which the clusters belong. Figure 11 The left subgraph in the middle represents the teacher target vector y extracted using the target pre-trained model. [patchs] It can cluster key parts of the human body, such as the face, feet, and knees. Figure 11 The right-hand subgraph represents the teacher's target vector y. [patchs]The ability to cluster pedestrians' necks, backpacks and their tracks, and bicycles demonstrates that the target pre-trained model has the ability to self-supervisedly extract key human body parts of pedestrians and fine-grained features of surrounding objects.

[0152] This automatic extraction capability is crucial for improving the accuracy of pedestrian re-identification. For example, automatic neck localization improves the recognition of accessories such as necklaces in pedestrian re-identification; automatic face and head localization improves hairstyle recognition; and automatic foot detection greatly enhances shoe recognition. Figure 12 As shown, even in complex scenes such as occlusion, incompleteness, and misaligned pedestrians, the pre-trained model can still extract pedestrian contours well while ignoring the influence of irrelevant background.

[0153] like Figure 13 As shown, the target pre-trained model can also capture the feature relationships between different images of the same pedestrian very well, even in cases with large deformations such as turning around or riding a bicycle.

[0154] Example 2:

[0155] like Figure 14 The diagram shown is a schematic representation of the architecture of a device for extracting pedestrian features according to an embodiment of the present invention. The device for extracting pedestrian features in this embodiment includes one or more processors 21 and a memory 22. Figure 14 Take a processor 21 as an example.

[0156] Processor 21 and memory 22 can be connected via a bus or other means. Figure 14 Taking the example of a connection between China and Israel via a bus.

[0157] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the method for extracting pedestrian features in this embodiment. The processor 21 executes the method for extracting pedestrian features by running the non-volatile software program and instructions stored in the memory 22.

[0158] Memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 22 may optionally include memory remotely located relative to processor 21, which can be connected to processor 21 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0159] The program instructions / modules are stored in the memory 22. When executed by one or more processors 21, they perform the method for extracting pedestrian features as described in the above embodiments, for example, performing the above-described... Figure 1 , Figure 3 , Figure 6 and Figures 8-10 The steps shown.

[0160] This invention also provides a non-volatile computer storage medium storing computer-executable instructions that are executed by one or more processors, for example... Figure 12 A processor 21 may enable one or more of the aforementioned processors to execute the method for extracting pedestrian features in a specific embodiment of the present invention, for example, to perform the above-described method. Figure 1 , Figure 3 , Figure 6 and Figures 8-10 The steps shown can also be implemented. Figure 14 The various modules and units described above; or the method for extracting pedestrian features as described in the specific embodiments of the present invention, for example, performing the above-described... Figure 1 , Figure 3 , Figure 6 and Figures 8-10 The steps shown can also be implemented. Figure 14 The various modules and units mentioned above.

[0161] It is worth noting that the information interaction and execution process between the modules and units in the above-mentioned device and system are based on the same concept as the processing method embodiment of the present invention. For details, please refer to the description in the method embodiment of the present invention, and will not be repeated here.

[0162] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0163] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for extracting pedestrian features, characterized in that, include: Unlabeled pedestrian images are input into the initial pre-trained model, and the unlabeled pedestrian images are converted into a sequence of original image blocks; The teacher network of the initial pre-trained model is used to extract pedestrian features from the original image patch sequence to obtain the teacher target vector; Randomly mask the original words in the original image block sequence to obtain a masked image block sequence; The student network using the initial pre-trained model extracts pedestrian features from the masked image block sequence to obtain a student target vector; wherein, it further includes: the original word unit includes a classification original word unit and at least one input original word unit; the classification original word unit is used as a classification mask word unit; for the at least one input original word unit, the corresponding mask amount is set to a first preset value to select to mask the input original word unit; the corresponding mask amount is set to a second preset value to select not to mask the input original word unit; when the mask amount of the input original word unit is the first preset value, the input original word unit is replaced with a mask to obtain a masked input original word unit; when the mask amount of the input original word unit is the second preset value, the input original word unit is used as an unmasked input original word unit; a masked image block sequence is obtained according to the classification mask word unit and the input image block sequence; Based on the teacher target vector and the student target vector, the network parameters of the initial pre-trained model are optimized; wherein, the optimization further includes: the teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector; the student target vector includes a representative student target vector and at least one ordinary student target vector; the pre-training loss function includes a mask reconstruction loss function and a global loss function; the mask reconstruction loss function is used to compare the at least one ordinary teacher target vector with the at least one ordinary student target vector to obtain a mask reconstruction loss value; the global loss function is used to compare the representative teacher target vector with the representative student target vector to obtain a global loss value. The above process is iterated until the change in the loss function of the initial pre-trained model is less than a preset convergence threshold, thus obtaining the target pre-trained model, which is then used to extract pedestrian features.

2. The method for extracting pedestrian features according to claim 1, characterized in that, The expression for the masked image block sequence is: ; in, For classification mask words, Given a sequence of image patches as input, The i-th input primitive word in the input image block sequence; , For masking, This is the mask value.

3. The method for extracting pedestrian features according to claim 1, characterized in that, The student target vector includes a representative student target vector and at least one ordinary student target vector; The student network using the initial pre-trained model extracts pedestrian features from the masked image block sequence to obtain the student target vector, which includes: The student network is used to extract local pedestrian features from the classification mask words in the masked image block sequence to obtain the first representative feature; the first representative feature is mapped to the target vector space to obtain the representative student target vector; The student network is used to predict the local pedestrian features of the masked original input words in the masked image block sequence, and the local pedestrian features of the unmasked original input words are extracted to obtain the second representative features; the second representative features are mapped to the target vector space to obtain the ordinary student target vector.

4. The method for extracting pedestrian features according to claim 1, characterized in that, The step of optimizing the network parameters of the initial pre-trained model based on the teacher target vector and the student target vector further includes: The product of the mask reconstruction loss value and the corresponding weight value is determined as the first intermediate value; the product of the global loss value and the corresponding weight value is determined as the second intermediate value; the sum of the first intermediate value and the second intermediate value is determined as the pre-training loss value, so as to calculate the pre-training loss value of this iteration through the pre-training loss function; Based on the pre-training loss value, the network parameters of the initial pre-trained model are updated, and the next iteration is performed according to the updated network parameters.

5. The method for extracting pedestrian features according to claim 1, characterized in that, The expression for the pre-training loss function is: ;in, The global loss function is... For the corresponding weight values; The loss function for mask reconstruction is... For the corresponding weight values; The expression for the global loss function is: ;in, To represent the teacher's target vector, To represent the student's target vector, The default loss function is used. The expression for the mask reconstruction loss function is: Among them, for n target vectors of ordinary teachers and n target vectors of ordinary students, Let i be the target vector of the i-th ordinary teacher. Let i be the target vector of the i-th ordinary student. This is the mask value.

6. The method for extracting pedestrian features according to claim 1, characterized in that, The original word unit includes a classification original word unit and at least one input original word unit; the teacher target vector includes a representative teacher target vector and at least one ordinary teacher target vector. The unlabeled pedestrian images are input into the initial pre-trained model, and the unlabeled pedestrian images are converted into a sequence of original image blocks; The teacher network, using the initial pre-trained model, extracts pedestrian features from the original image patch sequence to obtain the teacher target vector, which includes: The unlabeled pedestrian image is segmented into multiple image blocks; each image block is mapped to a patch word; from the patch words corresponding to the multiple image blocks, one patch word is selected as the original classification word, and the remaining patch words are used as the input original words to obtain the original image block sequence. The teacher network is used to extract global pedestrian features from the original word units classified in the original image block sequence to obtain the third representative feature; the third representative feature is mapped to the target vector space to obtain the representative teacher target vector. The teacher network is used to extract global pedestrian features from the input original words in the original image block sequence to obtain the fourth representative feature; the fourth representative feature is mapped to the target vector space to obtain the ordinary teacher target vector.

7. A device for extracting pedestrian features, characterized in that, The apparatus for extracting pedestrian features includes at least one processor and a memory, which are connected via a data bus. The memory stores instructions that can be executed by the at least one processor. After being executed by the processor, the instructions are used to implement the method for extracting pedestrian features according to any one of claims 1-6.

8. A non-volatile computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which are executed by one or more processors to perform the method for extracting pedestrian features as described in any one of claims 1-6.