A pedestrian re-identification method, device and medium

By introducing a two-stage dynamic search mechanism and early retreat strategy in pedestrian re-identification, the problem of improper allocation of computing resources in the existing technology is solved, efficient pedestrian re-identification is achieved, retrieval accuracy is ensured and calculation costs are reduced.

CN119516617BActive Publication Date: 2025-05-06CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510072083.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-06
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

When existing pedestrian re-identification methods handle queries of different difficulty levels, the allocation of computing resources is improperly, resulting in insufficient accuracy or waste of computing resources.

Method used

A two-stage dynamic retrieval mechanism is proposed to judge the difficulty of querying images through early withdrawal strategy and decide whether to enter the detailed reasoning stage. For simple queries, only global feature extraction is performed in the coarse inference stage to avoid unnecessary calculations; for difficult queries, fine-grained partial features are extracted in the fine inference stage.

Benefits of technology

While ensuring the accuracy of search, it effectively reduces computing costs, improves search efficiency, and realizes on-demand allocation of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516617B_ABST
    Figure CN119516617B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of surveillance video retrieval, and specifically relates to a pedestrian re-identification method, device and medium. The method focuses on balancing performance and efficiency, and proposes a two-stage dynamic retrieval mechanism, including a coarse reasoning stage and a fine reasoning stage. The early exit strategy is used to judge the difficulty of the query image to achieve on-demand allocation of computing resources: for simple queries, only global features are extracted in the coarse reasoning stage for rapid retrieval, and subsequent reasoning is terminated; for difficult queries, they are sent to the fine reasoning stage to further extract fine-grained partial features for refined retrieval. In addition, in order to further reduce the computing cost, the method of the present invention also introduces a hybrid expert model to dynamically activate the network module to participate in the calculation. By applying the method of the present invention, the computing cost can be effectively reduced while ensuring the retrieval accuracy, and the retrieval efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of surveillance video retrieval, and in particular relates to a pedestrian re-identification method, device and medium. Background Art

[0002] Person re-identification aims to retrieve the target pedestrian that matches a given query image from a large library of pedestrians captured by multiple different cameras. With the growing demand for public safety, person re-identification has been widely used in many fields such as video surveillance, criminal investigation, and smart city construction.

[0003] At present, most existing person re-ID methods focus on improving retrieval accuracy, often ignoring the importance of model computational efficiency. Specifically, these methods usually do not distinguish the difficulty of queries during inference, but use the same network processing for all queries and use unified features for retrieval. This practice may lead to insufficient accuracy when processing difficult queries, or waste unnecessary computing resources when processing simple queries. We experimentally found that the retrieval difficulty of different queries is often different: some query images have obvious discriminative visual features, and global features can obtain more accurate retrieval results. For these simple queries, over-reliance on fine-grained features may lead to false matches, because different pedestrians may have very similar appearances in some body parts, and not all part features are discriminative. On the contrary, for some difficult queries, such as pedestrian images that are difficult to identify due to occlusion, posture changes, or small inter-class differences, fine-grained partial features are required to support more refined matching.

[0004] In addition, existing pedestrian re-identification methods still have certain limitations when extracting fine-grained partial features: some methods extract partial features by introducing external models such as pose estimation or human body parsing models, but the performance of such methods is highly dependent on the accuracy of external models and increases the computational cost of the reasoning stage; other methods divide spatially adjacent image blocks or pixels into a group to learn part features, but do not fully consider the prior knowledge of human body topology, resulting in inaccurate partial positioning.

[0005] In summary, there is an urgent need for a pedestrian re-identification method, device and medium to solve the problems existing in the prior art. Summary of the invention

[0006] The present invention aims to provide a pedestrian re-identification method, device and medium, and the specific technical solution is as follows:

[0007] A pedestrian re-identification method comprises the following steps:

[0008] Construct a pedestrian re-identification algorithm for realizing pedestrian image re-identification, including a coarse reasoning stage, an early exit strategy and a fine reasoning stage. The coarse reasoning stage and the fine reasoning stage are respectively used to extract features of different granularities, and the early exit strategy is used to judge the difficulty of the query image, thereby determining whether the query image enters the fine reasoning stage;

[0009] Training the person re-identification algorithm, jointly training the coarse reasoning stage and the fine reasoning stage to obtain a person re-identification model;

[0010] The pedestrian image to be queried is input into the pedestrian re-identification model and matched with the images in the gallery to retrieve the target pedestrian in the gallery.

[0011] Optionally, the coarse reasoning stage includes coarse-grained feature extraction, and the process is as follows:

[0012] Split the pedestrian image into multiple blocks of fixed size;

[0013] Flatten all blocks and map them through a linear projection function to get block embedding;

[0014] Add learnable tags, position encoding embeddings, and camera embeddings to the front end of each block embedding to get the input sequence;

[0015] Call the encoder to process the input sequence and obtain global features and block features.

[0016] Optionally, the process of the early exit strategy is as follows:

[0017] Calculate the cosine similarity between the global feature and each block feature to obtain a set of similarity scores;

[0018] The similarity scores are sorted in descending order, and the first-order differences of the similarity scores are calculated to obtain a set of first-order difference values ​​of similarity;

[0019] The subscript index of the maximum value in the first-order difference value of similarity is used as the segmentation point. The block features before the segmentation point are regarded as the body area, and the block features after the segmentation point are regarded as the background or occlusion area.

[0020] An early exit threshold is set to determine whether the number of blocks belonging to the body area in the current query image can support high-confidence retrieval. If the current query image contains a sufficient number of visible body areas, it is considered a "simple" query and the subsequent reasoning is terminated early without entering the computationally intensive fine reasoning stage; otherwise, it is a "difficult" query and needs to be sent to the fine reasoning stage to extract fine-grained partial features for more refined retrieval.

[0021] Optionally, the fine reasoning stage follows the paradigm of a hybrid expert model, consisting of a block-partial routing network and a set of partial expert modules;

[0022] Block - Partially routed network:

[0023] Calculate the probability that each block belongs to the background or each body part category, and fuse the probabilities of each body part category to get the pedestrian foreground probability;

[0024] A soft routing mechanism is adopted to perform probability weighted average pooling on the block features according to the obtained probability distribution, and a pedestrian foreground feature, a background feature, and multiple body part features are obtained;

[0025] Some expert modules:

[0026] One-dimensional convolution is used to interact information between different feature channels for body part features to obtain the first part of features;

[0027] Use the multi-head cross attention mechanism for further processing, derive the query matrix from the first part of features, and derive the key matrix and value matrix from the foreground features. Calculate the multi-head cross attention based on the query matrix, key matrix and value matrix;

[0028] The fully connected layer and layer normalization are called to obtain new part features, which have the same dimension as the body part features.

[0029] In order to further reduce the computational cost and deal with background / occlusion noise more effectively, a binary partial visibility routing weight is generated based on the probability distribution output by the block-part routing network during inference. When the visibility routing weight of a body part is 1, the corresponding part expert module is activated during inference; otherwise, the corresponding part expert module is not activated during inference.

[0030] Optionally, during the training of the pedestrian re-identification algorithm, all training images go through a coarse reasoning stage and a fine reasoning stage to achieve joint training of the coarse reasoning stage and the fine reasoning stage, and all partial expert modules in the fine reasoning stage are activated and participate in the training.

[0031] Optionally, the total training loss includes the coarse inference phase loss, the routing loss in the fine inference phase, and some expert losses. The total loss The expression is as follows:

[0032] ;

[0033] ;

[0034] ;

[0035] ;

[0036] in, represents the coarse inference stage loss, Indicates the routing loss, Indicates expert loss; Indicates ID loss, Represents global features; represents the cross entropy loss, Represents the foreground features, Represents splicing The features obtained after the body part features, Respectively body part characteristics, represents the cross entropy loss with label smoothing, express The weight parameter, represents the separation loss function; Indicates New part features, Respectively New part features, represents the triplet loss.

[0037] Optionally, a two-stage dynamic retrieval mechanism dynamically selects the granularity of retrieval features and adjusts the network processing flow according to the difficulty of the query image, including a coarse reasoning stage and a fine reasoning stage.

[0038] Optionally, for the rough inference stage, the Euclidean distance of the global features is directly used to measure the similarity between the query and the gallery samples.

[0039] Optionally, for the detailed reasoning stage, a part-part matching strategy is used to measure the similarity between the query and the gallery sample. The process of the part-part matching strategy is as follows:

[0040] Calculate the distance of the foreground feature and the distance between all body part features calculated using the visibility routing weight values ​​of the body parts to obtain a first distance;

[0041] The visibility routing weight value of the body part is used to calculate the distance between all new part features to obtain a second distance;

[0042] In the fine reasoning stage, the distance between the query and the gallery sample is the sum of the first distance and the second distance.

[0043] Additionally, a computer device includes a memory and a processor;

[0044] The memory is used to store a computer program executable on the processor;

[0045] The processor is used to implement the steps of the above-mentioned pedestrian re-identification method when executing the computer program.

[0046] In addition, a computer-readable storage medium is provided, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-mentioned pedestrian re-identification method are implemented.

[0047] The application of the technical solution of the present invention has the following beneficial effects:

[0048] The present invention proposes a person re-identification method, which aims to balance model performance and reasoning efficiency. The method of the present invention proposes a novel two-stage dynamic retrieval mechanism. The mechanism includes a coarse reasoning stage and a fine reasoning stage, and uses an early exit strategy to judge the difficulty of the query image to achieve on-demand allocation of computing resources: for "simple" queries, global features are extracted only in the coarse reasoning stage for rapid retrieval, and subsequent reasoning is terminated to avoid unnecessary computing overhead; for "difficult" queries, they are sent to the fine reasoning stage to further extract fine-grained partial features for refined retrieval. The method of the present invention effectively reduces computing costs and improves retrieval efficiency while ensuring accuracy.

[0049] The method of the present invention innovatively applies the hybrid expert model to partial feature extraction for pedestrian re-identification. By using the human topology prior knowledge to guide the block-part routing network, the method of the present invention can achieve accurate body part positioning without introducing additional computational overhead. Each body part is assigned a part expert module to customize the learning of the fine-grained features of the part, and the part expert module is selectively activated according to the visibility routing weight during the reasoning process, thereby further reducing the computational cost.

[0050] In addition to the above-described purposes, features and advantages, the present invention has other purposes, features and advantages. The present invention will be further described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0052] Figure 1 It is a flowchart of the steps of the pedestrian re-identification method in a preferred embodiment of the present invention.

[0053] Figure 2 It is a schematic diagram of the technical solution of a preferred embodiment of the present invention.

[0054] Figure 3 It is a schematic diagram of part of the expert modules of the preferred embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0056] like Figure 1 , Figure 2 As shown, this embodiment provides a pedestrian re-identification method, including the following steps:

[0057] S1: Construct a pedestrian re-identification algorithm for realizing pedestrian image re-identification, including a coarse reasoning stage, an early exit strategy and a fine reasoning stage. The coarse reasoning stage and the fine reasoning stage are respectively used to extract features of different granularities, and the early exit strategy is used to judge the difficulty of the query image, thereby determining whether the query image enters the fine reasoning stage.

[0058] Rough reasoning stage:

[0059] In this embodiment, the coarse reasoning stage includes coarse-grained feature extraction. In this embodiment, ViT pre-trained on ImageNet is used as a feature extractor. The process is as follows:

[0060] Given a pedestrian query image ( ,in, Represent the number of channels, height and width of the image respectively), Split into Blocks of fixed size;

[0061] Flatten all blocks and pass them through a trainable linear projection function Mapping is performed to obtain a set of block embeddings :

[0062] ;

[0063] Add a learnable Tag (the tag can be represented as ), learnable position encoding embedding and camera embedding, and get the input sequence , the expression is as follows:

[0064] ;

[0065] in, Representation position encoding embedding , Indicates camera embedding , represents the hyperparameter that balances the camera embedding weights.

[0066] Call the encoder to process the input sequence , the final output features include global features and Block Features .

[0067] Early exit strategy:

[0068] In order to achieve a trade-off between retrieval performance and computational efficiency, this embodiment introduces an early exit mechanism, which allows the model to terminate subsequent reasoning in advance when it determines that the credibility of the query retrieval results is high enough, thereby avoiding unnecessary consumption of computing resources. However, traditional early exit mechanisms are mostly used in classification tasks, and early exit indicators are usually based on classifier scores. However, in the task of person re-identification, due to the inconsistency of training and testing categories and the unknown number of categories during testing, the traditional early exit indicator based on classifier scores is no longer applicable. This embodiment designs an early exit strategy for pedestrian re-identification, which evaluates the difficulty of the query image based on the first-order difference of similarity, thereby deciding whether to exit early. The process of the early exit strategy is as follows:

[0069] This embodiment evaluates the difficulty of the query image by counting the number of blocks belonging to the body area. First, calculate the global feature And each block feature The cosine similarity between , and obtain a set of similarity scores , the calculation expression is as follows:

[0070] ;

[0071] Similarity score Sort in descending order to get the sorted similarity score , and calculate the similarity score The first-order difference of , and a set of similarity first-order difference values ​​are obtained :

[0072] ;

[0073] This embodiment assumes that the pedestrian's body region and the background / occlusion region belong to different categories in the feature space. When the block features corresponding to the body region are transformed into the block features corresponding to the background / occlusion region, an obvious feature conversion should be shown. The subscript index of the maximum value in is used as the split point , before the split point The blocks are considered as body areas, and the remaining blocks after the segmentation point are considered as background or occluded areas;

[0074] Set an early exit threshold , a threshold is used to determine whether the number of blocks belonging to the body area contained in the current query image can support high-confidence retrieval. When , it indicates that the current query image contains enough visible body areas to support high-confidence retrieval, and is regarded as a simple query. The global features extracted in the coarse reasoning stage are directly used for retrieval, and the reasoning is terminated early without entering the computationally intensive fine reasoning stage. Otherwise, it is a "difficult" query and needs to be sent to the fine reasoning stage to extract fine-grained partial features for more refined retrieval.

[0075] Detailed reasoning stage:

[0076] The detailed reasoning stage follows the paradigm of the hybrid expert model, which consists of a block-part routing network and a set of part expert modules. By combining identity tags and human topology prior knowledge, the block-part routing network is responsible for distributing block features to the body parts to which they belong, thus achieving block-level body part localization. For each body part, a dedicated expert module is assigned to it to customize the learning of the fine-grained features of the part.

[0077] Block - Partially routed network:

[0078] In this embodiment, the block-part routing network is responsible for Block Features Route to the corresponding Part categories ,in Represents background, represent This embodiment uses a typical Router structure in the hybrid expert model to implement a block-part routing network, which consists of a parameter A fully connected layer and a Through the block-part routing network, the probability distribution of each block belonging to the background or each body part category can be calculated. , the calculation expression is as follows:

[0079] ;

[0080] in, express Blocks belong to some categories probability.

[0081] Then, the block-part routing network is based on the obtained probability distribution Distribute block features to some expert modules. The routing network in the existing hybrid expert model usually adopts the Top-k routing mechanism, that is, assigning a fixed number of However, due to the large differences in the area of ​​each body part, the Top-k routing mechanism is not suitable for pedestrian re-identification. If the setting is too small, it may not be enough to cover a large area of ​​the body (such as the torso), resulting in loss of information; If the setting is too large, a small body part (such as the head) may be assigned to the surrounding background / occlusion blocks, resulting in the introduction of noise. Therefore, this embodiment adopts a soft routing mechanism and does not require a fixed number of block features to be assigned to each body part expert. First, the fusion The probability of body part categories , get the pedestrian foreground probability , the expression is as follows:

[0082] ;

[0083] in, represents the summation calculation, Respectively represent Block features belong to body part categories probability.

[0084] use Block Features Perform probability weighted average pooling to obtain a pedestrian foreground feature , a background feature and Body part characteristics , the calculation formula is as follows:

[0085] ;

[0086] in, represents the first Block features.

[0087] Some expert modules:

[0088] In order to learn the details of each body part and make up for the deficiency of traditional Transformer in detail learning, this embodiment allocates a dedicated part expert module to each body part to focus on learning the fine-grained discriminative features of the part. As a coarse-grained prompt, it provides contextual support for the learning of fine-grained partial features to guide the learning of each expert module. The design scheme of the partial expert module is as follows: Figure 3 shown.

[0089] Characteristics of body parts Use one-dimensional convolution to interact and integrate information between different feature channels to obtain the first part of the features , the expression is as follows:

[0090] ;

[0091] in, Represents a one-dimensional convolution.

[0092] Use the multi-head cross attention mechanism for further processing. Export query matrix , by the foreground features Export bond matrix Sum Matrix , the expression is as follows:

[0093] ;

[0094] in, , , Represent the weight matrices of query, key, and value respectively.

[0095] According to the query matrix , the bond matrix Sum Matrix , calculate the multi-head cross attention, the calculation expression is as follows:

[0096] ;

[0097] in, is the number of attention heads, is the output transformation matrix integrating the outputs of all attention heads. The calculation expression is as follows:

[0098] ;

[0099] ;

[0100] in, Respectively represent The query matrix, key matrix, and value matrix of the attention head, is the scaling factor.

[0101] Call the fully connected layer and layer normalization to get new partial features , the new part features have the same dimension as the body part features.

[0102] The complete forward propagation calculation process of some expert modules is:

[0103] ;

[0104] ;

[0105] in, and represents two fully connected layers, and Representation layer normalization operation, It is a body part feature New partial features output by the partial expert module.

[0106] In order to further reduce the computational cost and more effectively deal with background or occlusion noise, this embodiment dynamically selects some expert modules to activate and participate in the calculation according to the visibility of the body parts during reasoning. This embodiment is based on the probability distribution of the block-part routing network output , generating binary partial visibility routing weights Specifically, for body parts ,if only The probability value of at least one block in , then the partial visibility routing weight of the body part , activate the corresponding expert modules during reasoning; otherwise , some expert modules are not activated during reasoning:

[0107] ;

[0108] in, express The probability value of at least one block in .

[0109] S2: training the person re-identification algorithm, performing joint training on the coarse reasoning stage and the fine reasoning stage, and obtaining a person re-identification model.

[0110] During the training of the person re-identification algorithm, all training images go through a coarse reasoning stage and a fine reasoning stage, achieving joint training of the coarse reasoning stage and the fine reasoning stage, and all partial expert modules in the fine reasoning stage are activated and participate in the training.

[0111] Optionally, the total loss of training includes the coarse inference stage loss as well as the routing loss and part of the expert loss in the fine inference stage. The expression is as follows:

[0112] ;

[0113] in, represents the coarse inference stage loss, Indicates the routing loss, Indicates expert loss;

[0114] Coarse inference stage loss Use cross entropy loss as ID loss to supervise the global features learned by the encoder , the expression is as follows:

[0115] ;

[0116] in, Indicates ID loss, Represents global features.

[0117] Routing loss Responsible for supervising the training of the block-part routing network, the expression is as follows:

[0118] ;

[0119] ;

[0120] in, represents the cross entropy loss, Represents the foreground features, Represents splicing The features obtained after the body part features, Respectively body part characteristics, represents the cross entropy loss with label smoothing, express The weight parameter, represents the separation loss function.

[0121] Cross Entropy Loss with Label Smoothing Use rough body part labels to introduce human topology prior knowledge to guide the routing of block features to body parts. , some of its labels , where 0 is the background label and 1 to yes The label of a body part. The expression is as follows:

[0122] ;

[0123] ;

[0124] in, represents the batch size, represents the label smoothing regularization rate, Indicates Block features belong to body part categories probability.

[0125] Separation loss It is used to separate the body area from the background, occlusion and other noise areas, so that the block-routing network is more focused on the pedestrian's body area and realizes the recognition of the body area. The expression is as follows:

[0126] ;

[0127] in, Indicates the first image in the batch The background features of the pedestrian image, is the first image in the batch Foreground features of a pedestrian image.

[0128] Some experts lost Responsible for supervising the learning of fine-grained features by some expert modules, the expression is as follows:

[0129] ;

[0130] in, Indicates ID loss, Indicates New part features, Respectively New part features, represents the triplet loss.

[0131] S3: Input the pedestrian image to be queried into the pedestrian re-identification model and match it with the images in the gallery to retrieve the target pedestrian in the gallery.

[0132] The two-stage dynamic retrieval mechanism proposed in this embodiment for pedestrian re-identification dynamically selects the granularity of retrieval features and adjusts the network processing flow according to the difficulty of the query image, including a coarse reasoning stage and a fine reasoning stage.

[0133] For the rough reasoning stage, the Euclidean distance of the global features is directly used to measure the query and gallery samples Similarity:

[0134] ;

[0135] in, is the global feature output by the encoder in the coarse inference stage, is the Euclidean distance metric.

[0136] For the fine reasoning stage, a part-part matching strategy is used to weigh the query and gallery samples The process of the part-to-part matching strategy is as follows:

[0137] Calculate the distance of the foreground feature and the distance between all body part features calculated using the visibility routing weight values ​​of the body parts to obtain a first distance;

[0138] The visibility routing weight value of the body part is used to calculate the distance between all new part features to obtain a second distance;

[0139] In the detailed reasoning phase, the query and gallery samples The distance between is the sum of the first distance and the second distance, and the calculation expression is as follows:

[0140] ;

[0141] ;

[0142] ;

[0143] in, represents the first distance, represents the second distance, and Both indicate The visibility routing weight values ​​of the body parts are used to ensure that only co-activated body parts are compared. If there are no co-activated body parts between two samples, their distance is set to infinity.

[0144] To verify the effectiveness of the method in this embodiment, the retrieval accuracy evaluation indicators recognized in the field of pedestrian re-identification, Rank-1 and mAP, are compared with the most advanced methods on the occluded pedestrian dataset Occluded-Duke and two complete pedestrian datasets Market-1501 and DukeMTMC-ReID, as shown in Table 1. Among them, PGFA, HOReID, IGOAS, BPBReID, RTGAT, and GPEOG are methods based on convolutional neural networks (CNN), and PAT, TransReID, FED, DRL-Net, and SCAT are methods based on Transformer. In addition, PGFA, HOReID, GPEOG, and PFD introduce external models such as posture estimation or human body analysis for partial feature extraction, and PAT and DRL-Net add CNN to the Transformer structure. It can be found from Table 1 that the method of this embodiment surpasses all the compared methods in both simple scenes and highly challenging occlusion scenes, and obtains the highest Rank-1 and mAP, demonstrating that the method of this embodiment has good versatility and stability in various data sets and scenes, indicating its robustness and superiority.

[0145] Table 1 Comparison results of the retrieval accuracy index Rank-1 and mAP of the method in this embodiment and the current most advanced method on the occluded pedestrian dataset and the complete pedestrian dataset

[0146]

[0147] In addition, to verify the trade-off between retrieval performance and computational efficiency of this example method, experiments are conducted on the Occluded-Duke dataset, using Rank-1 accuracy and mAP as performance indicators, and FLOPs in the fine reasoning stage as efficiency indicators. The results are shown in Table 2. Early exit threshold It is an important factor affecting the balance between retrieval performance and computational efficiency in this example method, because this example method uses an early exit threshold To evaluate the difficulty of the query and control the number of queries entering the detailed reasoning stage. This causes more queries to be identified as “difficult”, which in turn causes more queries to enter the computationally intensive fine reasoning stage. As can be seen from Table 2, as As the rank-1 accuracy increases, FLOPs decreases. This is because more and more queries are judged as "simple" queries, resulting in fewer and fewer "difficult" queries entering the detailed reasoning stage, reducing the computational cost of reasoning. The decrease of increases first and then decreases: the increase first may be because for some queries, not all local features are discriminative, that is, two images with different identities may be very similar in appearance in some parts, resulting in mismatches. On the contrary, using only the global features in the coarse reasoning stage can avoid this situation; the decrease may be because the queries sent to the fine reasoning stage do not fully cover the difficult queries that require fine-grained partial features to assist in retrieval. For mAP, it increases with The decrease of shows an overall downward trend, which indicates that using only global features is not a good way to retrieve, and it is also important to use fine-grained partial features for detail matching. The above results show that compared with the traditional unified retrieval mechanism that does not distinguish the difficulty of the query, the two-stage dynamic retrieval mechanism proposed in this example method can achieve a trade-off between model performance and model efficiency by adjusting the confidence. While bringing performance gains, it can also reduce unnecessary computing costs and speed up reasoning, and realize the adaptive allocation of computing resources between "simple" and "difficult" samples.

[0148] Table 2 Effects of different early exit thresholds on the performance and efficiency of CFPER

[0149]

[0150] In addition, this embodiment also provides a computer device, including a memory and a processor;

[0151] The memory is used to store a computer program executable on the processor;

[0152] The processor is used to implement the steps of the above-mentioned pedestrian re-identification method when executing the computer program.

[0153] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the computer device.

[0154] The computer device may be a computing device such as a mobile phone, a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may include, but is not limited to, a processor and a memory. For example, the computer device may also include an input / output device, a network access device, a bus, etc.

[0155] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer device, and uses various interfaces and lines to connect various parts of the entire computer device.

[0156] The memory can be used to store the computer program and / or module, and the processor implements the computer program by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0157] Wherein, if the module / unit integrated in the computer device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.

[0158] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned pedestrian re-identification method are implemented.

[0159] This embodiment proposes a pedestrian re-identification method, which aims to balance model performance and reasoning efficiency. This embodiment method proposes a novel two-stage dynamic retrieval mechanism. The mechanism includes a coarse reasoning stage and a fine reasoning stage, and uses an early exit strategy to judge the difficulty of the query image to achieve on-demand allocation of computing resources: for "simple" queries, only global features are extracted for fast retrieval in the coarse reasoning stage, and subsequent reasoning is terminated to avoid unnecessary computing overhead; for "difficult" queries, they are sent to the fine reasoning stage to further extract fine-grained partial features for refined retrieval. The method of the present invention effectively reduces computing costs and improves retrieval efficiency while ensuring accuracy.

[0160] The method of this embodiment innovatively applies the hybrid expert model to partial feature extraction for pedestrian re-identification. By using the human topology prior knowledge to guide the block-part routing network, the method of the present invention can achieve accurate body part positioning without introducing additional computational overhead. Each body part is assigned a part expert module to customize the learning of the fine-grained features of the part, and the part expert module is selectively activated according to the visibility routing weight during the reasoning process, thereby further reducing the computational cost.

[0161] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0162] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A pedestrian re-identification method, characterized in that: The steps include: Constructing a pedestrian re-identification algorithm for realizing pedestrian image re-identification, including a coarse reasoning stage, an early exit strategy and a fine reasoning stage, wherein the coarse reasoning stage and the fine reasoning stage are respectively used to extract features of different granularities, and the early exit strategy is used to judge the difficulty of the query image, thereby determining whether the query image enters the fine reasoning stage; Training the person re-identification algorithm, jointly training the coarse reasoning stage and the fine reasoning stage to obtain a person re-identification model; Input the pedestrian image to be queried into the pedestrian re-identification model and match it with the images in the gallery to retrieve the target pedestrian in the gallery; The process of the early exit strategy is as follows: Calculate the cosine similarity between the global feature and each block feature to obtain a set of similarity scores; The similarity scores are sorted in descending order, and the first-order differences of the similarity scores are calculated to obtain a set of first-order difference values ​​of similarity; The subscript index of the maximum value in the first-order difference value of similarity is used as the segmentation point. The block features before the segmentation point are regarded as the body area, and the block features after the segmentation point are regarded as the background or occlusion area. An early exit threshold is set to determine whether the number of blocks belonging to the body area contained in the current query image can support high-confidence retrieval. If the current query image contains a sufficient number of visible body areas, the current query image is regarded as a simple query, and subsequent reasoning is terminated in advance without entering the computationally intensive detailed reasoning stage. Otherwise, the current query image is a difficult query, and the current query image is sent to the fine reasoning stage to extract fine-grained partial features and perform subsequent reasoning.

2. The pedestrian re-identification method according to claim 1, characterized in that: The coarse reasoning stage includes coarse-grained feature extraction, and the process is as follows: Split the pedestrian image into multiple blocks of fixed size; Flatten all blocks and map them through a linear projection function to get block embedding; Add learnable tags, position encoding embeddings, and camera embeddings to the front end of each block embedding to get the input sequence; Call the encoder to process the input sequence and obtain global features and block features.

3. The pedestrian re-identification method according to claim 1, characterized in that: The detailed reasoning stage follows the paradigm of hybrid expert model, including a block-partial routing network and a set of partial expert modules; Block - Partially routed network: Calculate the probability that each block belongs to the background or each body part category, and fuse the probabilities of each body part category to get the pedestrian foreground probability; A soft routing mechanism is adopted to perform probability weighted average pooling on the block features according to the obtained probability distribution, and a pedestrian foreground feature, a background feature, and multiple body part features are obtained; Some expert modules: One-dimensional convolution is used to interact information between different feature channels for body part features to obtain the first part of features; Use the multi-head cross attention mechanism for further processing, derive the query matrix from the first part of features, derive the key matrix and value matrix from the foreground features, and calculate the multi-head cross attention based on the query matrix, key matrix and value matrix; Calling a fully connected layer and layer normalization to obtain a new part feature, wherein the new part feature has the same dimension as the body part feature; During reasoning, a binary partial visibility routing weight is generated based on the probability distribution output by the block-part routing network. When the visibility routing weight of a body part is 1, the corresponding partial expert module is activated during reasoning. Otherwise, the corresponding expert modules will not be activated during reasoning.

4. The pedestrian re-identification method according to claim 1, characterized in that: During the training of the person re-identification algorithm, all training images go through a coarse reasoning stage and a fine reasoning stage, achieving joint training of the coarse reasoning stage and the fine reasoning stage, and all partial expert modules in the fine reasoning stage are activated and participate in the training.

5. The pedestrian re-identification method according to claim 1, characterized in that: The total loss of training includes the loss of the coarse reasoning stage, the routing loss of the fine reasoning stage, and the partial expert loss. The expression is as follows: in, represents the coarse inference stage loss, Indicates the routing loss, Indicates expert loss; represents ID loss, G represents global features; represents the cross entropy loss, r f represents the foreground feature, r c represents the feature obtained by concatenating M body part features, r1,…,r M Represent M body part features respectively, represents the cross entropy loss with label smoothing, λ h express The weight parameter, represents the separation loss function; e i represents the i-th new partial feature, e1,…,e M Represent M new partial features respectively, represents the triplet loss.

6. The pedestrian re-identification method according to claim 1, characterized in that: The two-stage dynamic retrieval mechanism dynamically selects the granularity of retrieval features and adjusts the network processing flow according to the difficulty of the query image, including the coarse reasoning stage and the fine reasoning stage.

7. The pedestrian re-identification method according to claim 6, characterized in that: For the rough reasoning stage, the Euclidean distance of global features is directly used to measure the similarity between the query and the gallery samples; For the detailed reasoning stage, a partial-to-partial matching strategy is used to measure the similarity between the query and the gallery samples. The process of the partial-to-partial matching strategy is as follows: Calculate the distance of the foreground feature and the distance between all body part features calculated using the visibility routing weight values ​​of the body parts to obtain a first distance; The visibility routing weight value of the body part is used to calculate the distance between all new part features to obtain a second distance; In the fine reasoning stage, the distance between the query and the gallery sample is the sum of the first distance and the second distance.

8. A computer device, characterized in that: including memory and processor; The memory is used to store a computer program executable on the processor; The processor is configured to implement the steps of the pedestrian re-identification method according to any one of claims 1 to 7 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the pedestrian re-identification method according to any one of claims 1 to 7 are implemented.