Pedestrian re-identification model training method, recognition method, system, device and medium

By using the mask generator to extract local features and construct training data sets, the existing pedestrian recognition model has been solved, and the pedestrian recognition model with high accuracy and real-time detection and recognition is achieved.

CN113989838BActive Publication Date: 2025-05-13SUN YAT SEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111247638.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-26
Publication Date
2025-05-13
Estimated Expiration
2041-10-26

AI Technical Summary

Technical Problem

The existing pedestrian re-identification model has a low accuracy rate in the case of occlusion, with large parameters, long training and inference time, and the latest transformer-based model has a large amount of calculation and inference time, so it is impossible to realize real-time detection and recognition.

Method used

By obtaining the global feature information of the preset pedestrian image, and using the mask generator to generate the mask image for local feature extraction, the local feature vector and global feature vector are determined, the training data set is constructed and input into the convolutional neural network for training, a high-accuracy pedestrian re-identification model is obtained.

Benefits of technology

The accuracy of the pedestrian re-identification model in the occlusion situation is improved, the calculation amount and inference time of the model are reduced, real-time detection and recognition are realized, and the stability of the model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113989838B_ABST
    Figure CN113989838B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian re-identification model training method, recognition method, system, device and medium. The training method includes: obtaining a preset first pedestrian image, and determining the first global feature information and the second global feature information of the first pedestrian image; inputting the second global feature information into a preset mask generator to obtain multiple mask images, and then extracting local features from the first global feature information according to the multiple mask images to obtain multiple first local feature information; determining a local feature vector according to the multiple first local feature information, and determining a global feature vector according to the first global feature information; determining a training data set according to the local feature vector and the global feature vector, and inputting the training data set into a pre-constructed convolutional neural network for training to obtain a trained pedestrian re-identification model. The present invention improves the re-identification accuracy and stability of the pedestrian re-identification model, and can be widely used in the field of pedestrian re-identification technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and in particular to a pedestrian re-identification model training, identification method, system, device and medium. Background Art

[0002] Person re-identification is a technology that uses computer vision technology to associate pedestrians that appear in camera A with pedestrians that appear in camera B. The usual practice is to retrieve other classified pedestrian images given a search image to determine the identity of the pedestrian. As a subtask of cross-camera tracking, an efficient person re-identification model with hardware potential is a necessary condition for the implementation of the algorithm. In addition, since occlusions occur from time to time, the model also needs to maintain good performance in the presence of occlusions. Existing person re-identification models have the following problems:

[0003] (1) Some earlier networks that did not introduce additional information, such as VPM and PCB, had low re-identification accuracy under occlusion;

[0004] (2) Networks that introduce additional information such as pedestrian key points and pedestrian semantics for pedestrian re-identification under occlusion, such as PGFA and MaskReID, usually have a large number of parameters and take a long time to train and infer.

[0005] (3) Although the latest transformer-based model TransReID has good recognition performance, it has large computational complexity and long inference time, and cannot achieve real-time detection and recognition. Summary of the invention

[0006] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.

[0007] To this end, an object of an embodiment of the present invention is to provide a method for training a pedestrian re-identification model, which can improve the re-identification accuracy of the obtained pedestrian re-identification model.

[0008] Another object of an embodiment of the present invention is to provide a method for pedestrian re-identification with high accuracy.

[0009] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:

[0010] In a first aspect, an embodiment of the present invention provides a method for training a pedestrian re-identification model, comprising the following steps:

[0011] Acquire a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image;

[0012] Inputting the second global feature information into a preset mask generator to obtain a plurality of mask images, and then performing local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information;

[0013] Determine a local feature vector according to the first local feature information, and determine a global feature vector according to the first global feature information;

[0014] A training data set is determined according to the local feature vector and the global feature vector, and the training data set is input into a pre-built convolutional neural network for training to obtain a trained person re-identification model.

[0015] Further, in one embodiment of the present invention, the step of determining the first global feature information and the second global feature information of the first pedestrian image is specifically:

[0016] Extracting first global feature information and second global feature information of the first pedestrian image through a global feature extraction network;

[0017] The width of the first global feature information is the same as the width of the second global feature information, and the height of the first global feature information is the same as the height of the second global feature information.

[0018] Further, in one embodiment of the present invention, the step of extracting local features from the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information is specifically:

[0019] A tensor product operation is performed on the mask image and the first global feature information to obtain a plurality of first local feature information.

[0020] Further, in an embodiment of the present invention, the step of determining a local feature vector according to the plurality of first local feature information and determining a global feature vector according to the first global feature information specifically includes:

[0021] Performing global weighted pooling on the plurality of first local feature information respectively to obtain a plurality of second local feature information;

[0022] Determine a plurality of first feature vectors according to a plurality of the second local feature information, and then perform splicing processing on the plurality of the first feature vectors to obtain a local feature vector;

[0023] Perform global maximum pooling on the first global feature information to obtain third global feature information, and then determine a global feature vector according to the third global feature information.

[0024] Further, in one embodiment of the present invention, the step of determining the training data set according to the local feature vector and the global feature vector specifically includes:

[0025] Determine a training sample according to the local feature vector and the global feature vector;

[0026] Acquire annotation information of the first pedestrian image, and generate a classification label according to the annotation information;

[0027] A training data set is determined according to the training samples and corresponding classification labels.

[0028] Furthermore, in one embodiment of the present invention, the step of inputting the training data set into a pre-built convolutional neural network for training to obtain a trained person re-identification model specifically includes:

[0029] Inputting the training data set into the convolutional neural network to obtain a prediction classification result;

[0030] Determining a training loss value according to the predicted classification result and the classification label;

[0031] Updating the parameters of the convolutional neural network layer by layer through a back propagation algorithm according to the loss value;

[0032] When the loss value reaches a preset first threshold or the number of iterations reaches a preset second threshold, the training is stopped to obtain the pedestrian re-identification model.

[0033] In a second aspect, an embodiment of the present invention provides a method for pedestrian re-identification, comprising the following steps:

[0034] Obtain a second person image to be identified;

[0035] The second pedestrian image is input into the pedestrian re-identification model obtained by the pedestrian re-identification model training method as described in the first aspect to obtain a pedestrian re-identification result.

[0036] In a third aspect, an embodiment of the present invention provides a pedestrian re-identification model training system, including:

[0037] A global feature extraction module, used to obtain a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image;

[0038] A local feature extraction module, used for inputting the second global feature information into a preset mask generator to obtain a plurality of mask images, and then performing local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information;

[0039] a feature vector determination module, configured to determine a local feature vector according to the plurality of first local feature information, and determine a global feature vector according to the first global feature information;

[0040] The model training module is used to determine a training data set according to the local feature vector and the global feature vector, and input the training data set into a pre-built convolutional neural network for training to obtain a trained pedestrian re-identification model.

[0041] In a fourth aspect, an embodiment of the present invention provides a pedestrian re-identification model training device, comprising:

[0042] at least one processor;

[0043] at least one memory for storing at least one program;

[0044] When the at least one program is executed by the at least one processor, the at least one processor implements the pedestrian re-identification model training method described in the first aspect.

[0045] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and when the program executable by the processor is executed by the processor, it is used to execute the pedestrian re-identification model training method described in the first aspect or the pedestrian re-identification method described in the second aspect.

[0046] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:

[0047] The embodiment of the present invention obtains a preset first pedestrian image and determines the first global feature information and the second global feature information, then inputs the second global feature information into a preset mask generator to obtain multiple mask images, and then performs local feature extraction on the first global feature information according to the mask image to obtain multiple first local feature information, and then determines the local feature vector and the global feature vector according to the first local feature information and the first global feature information, so that the training data set for convolutional neural network training can be determined according to the local feature vector and the global feature vector, and the pedestrian re-identification model is obtained by training. The embodiment of the present invention can generate a semantically consistent mask image in a self-supervised manner through the mask generator, remove background noise, and improve the accuracy of pedestrian local feature extraction, thereby improving the re-identification accuracy of the pedestrian re-identification model; using the mask image to extract local features can avoid mutual influence between each local feature, further improve the re-identification accuracy of the pedestrian re-identification model, and also improve the stability of the pedestrian re-identification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0049] Figure 1 A flowchart of a method for training a pedestrian re-identification model provided by an embodiment of the present invention;

[0050] Figure 2 A schematic diagram showing the comparison of Rank1 and mAP of a person re-identification model obtained under different numbers of mask images provided by an embodiment of the present invention;

[0051] Figure 3 A schematic diagram showing the comparison of Rank1 and mAP of a person re-identification model obtained under different local-global feature ratios provided by an embodiment of the present invention;

[0052] Figure 4 A flowchart of a method for pedestrian re-identification provided by an embodiment of the present invention;

[0053] Figure 5 A structural block diagram of a pedestrian re-identification model training system provided by an embodiment of the present invention;

[0054] Figure 6 A structural block diagram of a pedestrian re-identification model training device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0055] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.

[0056] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.

[0057] Reference Figure 1 The embodiment of the present invention provides a method for training a pedestrian re-identification model, which specifically includes the following steps:

[0058] S101: Acquire a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image.

[0059] Specifically, a first pedestrian image can be obtained from the labeled pedestrian re-identification image set, and the first global feature information F of the first pedestrian image can be extracted using the global feature extraction network. g and the second global feature information

[0060] As an optional implementation, the step of determining the first global feature information and the second global feature information of the first pedestrian image is specifically as follows:

[0061] Extracting first global feature information and second global feature information of the first pedestrian image through a global feature extraction network;

[0062] The width of the first global feature information is the same as the width of the second global feature information, and the height of the first global feature information is the same as the height of the second global feature information.

[0063] Specifically, the first global feature information F g and the second global feature information They can be represented as (c, h, w) and (c*, h, w) respectively. They can be the same feature or different features. The only requirement is that their height h and width w must be equal. For features with fewer channels c*, for example, when using a ResNet-50 network and the final downsampling step is set to 1, the output of Conv4 can be used as The output of Conv5 is used as F g , at this time c*=1024, c=2048, which can not only reduce the amount of calculation but also improve the generalization ability.

[0064] S102: Input the second global feature information into a preset mask generator to obtain a plurality of mask images, and then perform local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information.

[0065] Specifically, for the features Apply a mask generator to generate p mask images M1, M2, ..., M p , and use these p mask images to identify the feature F g Processing generates p first local feature information F l,1, F l,2 , ..., F l,p In the embodiment of the present invention, the mask generator is not explicitly supervised, and the p mask images generated by the mask generator are all generated by self-supervision. Each mask image is actually equivalent to a local feature extraction mode.

[0066] As an optional implementation, the step of extracting local features from the first global feature information according to the multiple mask images to obtain multiple first local feature information is specifically as follows:

[0067] A tensor product operation is performed on the mask image and the first global feature information to obtain a plurality of first local feature information.

[0068] Specifically, the mask generator only requires a 1×1 convolution, feature After passing through the mask generator, p mask images M1, M2, ..., M are obtained. p , use these mask images to identify the features F g Extract local features. The local feature extraction formula is:

[0069] i=1,2,...,p

[0070] Thus, p first local feature information are obtained.

[0071] S103: Determine a local feature vector according to the plurality of first local feature information, and determine a global feature vector according to the first global feature information.

[0072] Specifically, the first local feature information and the first global feature information are converted into space-independent features, and the p first local feature information are respectively subjected to global weighted pooling, while the first global feature information is subjected to global maximum pooling to compress the spatial information of the feature into 1×1. Step S103 specifically includes the following steps:

[0073] S1031, performing global weighted pooling on the multiple first local feature information respectively to obtain multiple second local feature information;

[0074] S1032, determining a plurality of first feature vectors according to the plurality of second local feature information, and then concatenating the plurality of first feature vectors to obtain a local feature vector;

[0075] S1033: Perform global maximum pooling on the first global feature information to obtain third global feature information, and then determine a global feature vector according to the third global feature information.

[0076] Specifically, the embodiment of the present invention adopts a global weighted pooling operation on the first local feature information to obtain the space-independent second local feature information fl,i The formula for global weighted pooling is as follows:

[0077]

[0078] The global maximum pooling operation is applied to the first global feature information to obtain the third global feature information f that is independent of space. g After the fully connected layer, the second local feature information f l,i Converted into the first eigenvector of 256 dimensions, the third global feature information f g Converted to a 512-dimensional global feature vector, the ratio of local features to global features is controlled to be Finally, the first eigenvectors are concatenated to obtain the local eigenvector. During the training process, the classification loss and triplet loss of the local eigenvector and the global eigenvector are calculated respectively.

[0079] It can be appreciated that after the embodiment of the present invention uses the mask image to extract local features, the extracted features are globally weighted pooled according to the mask image, and different local features are processed independently before the final splicing to avoid mutual influence between local features.

[0080] S104, determining a training data set according to the local feature vector and the global feature vector, and inputting the training data set into a pre-built convolutional neural network for training to obtain a trained person re-identification model.

[0081] Specifically, the embodiment of the present invention uses the labeled first pedestrian image to determine the classification label, takes the local feature vector and the global feature vector obtained previously as training samples, and performs training with the goal of minimizing the loss between the predicted value and the true value, wherein the loss includes the triple loss of the features output after splicing and the classification loss.

[0082] As an optional implementation, the step of determining the training data set according to the local feature vector and the global feature vector specifically includes:

[0083] A1. Determine training samples based on local feature vectors and global feature vectors;

[0084] A2. Obtain the annotation information of the first person image and generate a classification label based on the annotation information;

[0085] A3. Determine the training data set based on the training samples and the corresponding classification labels.

[0086] As an optional implementation, the training data set is input into a pre-built convolutional neural network for training to obtain a trained person re-identification model, which specifically includes:

[0087] B1. Input the training data set into the convolutional neural network to obtain the predicted classification results;

[0088] B2. Determine the training loss value based on the predicted classification results and classification labels;

[0089] B3. Update the parameters of the convolutional neural network layer by layer through the back propagation algorithm according to the loss value;

[0090] B4. When the loss value reaches a preset first threshold or the number of iterations reaches a preset second threshold, the training is stopped to obtain a pedestrian re-identification model.

[0091] Specifically, the model training process is an alternating cycle of forward propagation and back propagation. Forward propagation extracts the input pedestrian image layer by layer, outputs the predicted classification result after a series of convolution and pooling operations, and calculates the loss with the actual classification label. After the loss value is calculated by the loss function, the convolution network parameters are updated layer by layer through the back propagation algorithm, thus completing an iterative training. When the loss function value reaches a given value or the number of iterations reaches the specified maximum value, the training is stopped and the pedestrian re-identification model is saved.

[0092] In the embodiment of the present invention, after the data in the training data set is input into the initialized convolutional neural network, the predicted classification result output by the model can be obtained, and the accuracy of the recognition model prediction can be evaluated according to the predicted classification result and the aforementioned classification label, so as to update the parameters of the model. For the pedestrian re-identification model, the accuracy of the model prediction result can be measured by the loss function, which is defined on a single training data and is used to measure the prediction error of a training data. Specifically, the loss value of the training data is determined by the label of a single training data and the prediction result of the model for the training data. In actual training, a training data set has a lot of training data, so the cost function is generally used to measure the overall error of the training data set. The cost function is defined on the entire training data set and is used to calculate the average value of the prediction error of all training data, which can better measure the prediction effect of the model. For a general machine learning model, based on the aforementioned cost function, plus the regularization term that measures the complexity of the model, it can be used as the objective function of the training, and the loss value of the entire training data set can be calculated based on the objective function. There are many types of commonly used loss functions, such as 0-1 loss function, square loss function, absolute loss function, logarithmic loss function, cross entropy loss function, etc., which can all be used as loss functions of machine learning models, which will not be elaborated here one by one. In an embodiment of the present invention, any loss function can be selected to determine the loss value of training. Based on the loss value of training, the back propagation algorithm is used to update the parameters of the model, and a trained point cloud tower recognition model can be obtained after several rounds of iteration. The specific number of iterations can be set in advance, or the training is considered to be completed when the test set meets the accuracy requirements.

[0093] The above describes the training steps of the pedestrian re-identification model of the embodiment of the present invention. It can be recognized that the embodiment of the present invention can self-supervise and generate a semantically consistent mask image through a mask generator, remove background noise, and improve the accuracy of local feature extraction of pedestrians, thereby improving the re-identification accuracy of the pedestrian re-identification model; using the mask image for local feature extraction can avoid the mutual influence between the local features, further improve the re-identification accuracy of the pedestrian re-identification model, and also improve the stability of the pedestrian re-identification model.

[0094] The pedestrian re-identification model of the embodiment of the present invention is further verified and explained through experiments below.

[0095] The embodiment of the present invention conducts an ablation experiment on the number of masks p, the local-global feature ratio γ, and the pooling of local features on the Occluded-Duke dataset, verifying the effectiveness of the self-supervised mask method, local feature weighted pooling, and local-global feature ratio proposed in the embodiment of the present invention. Figure 2The figure shows a comparison diagram of Rank1 (the probability that the top image in the search result is the correct result) and mAP (mean Average Precision, used to measure the search ability of the algorithm) of the pedestrian re-identification model obtained under different mask image numbers (Mask Number) p provided by the embodiment of the present invention. Compared with the traditional method of dividing local features into 6 parts, dividing local features into 18 parts through mask images obviously has better recognition performance, and the plateau period appears in the part from 6 to 12, which also explains why it has not been found that the performance of more divisions is better. Figure 3 The figure shows a comparison diagram of Rankl and mAP of the pedestrian re-identification model obtained under different local-global feature ratios γ provided by an embodiment of the present invention. In the process of increasing the local-global feature ratio from small to large, the model performance first increases and then decreases, and the model performance reaches the highest point when the local-global feature ratio is 9.

[0096] Table 1 below shows the Rank1 and mAP of the pedestrian re-identification model obtained by using different pooling methods for local features. It can be seen that the global weighted pooling (WP) proposed in the embodiment of the present invention significantly improves the recognition effect of the pedestrian re-identification model compared with the traditional global average pooling (GAP).

[0097] Pooling Rank1 mAP WP 62.2 52.3 GAP 59.2 50.3

[0098] Table 1

[0099] The Rank1 and mAP of the person re-identification model obtained by the embodiment of the present invention are compared with those of other similar models, as shown in Table 2 below. It can be found that the present invention is at the forefront in both Rank1 and mAP, which verifies the advanced nature of the method of the present invention.

[0100]

[0101] Table 2

[0102] In addition, the pedestrian remodel of the embodiment of the present invention has a computational load of 6.9G FLOPS and a parameter load of 56M, which is 68% more computational load and 14% more parameter load than the baseline model (PCB, 4.1G FLOPS, 49M).

[0103] Based on the above analysis and data set verification, it can be recognized that the self-supervised mask, local feature weighted pooling, and local-global feature ratio proposed in the embodiment of the present invention are effective in solving the problem of occluded pedestrian re-identification, which can greatly improve the accuracy and stability of occluded pedestrian re-identification, and reach the world's leading level in the main indicators Rank1 and mAP. On the other hand, the model structure proposed in the embodiment of the present invention does not introduce operations other than multiplication, division, addition, and absolute value calculation, which is conducive to the implementation of subsequent algorithms and has great reference significance for the design of future ReID models.

[0104] Reference Figure 4 The embodiment of the present invention provides a pedestrian re-identification method, which specifically includes the following steps:

[0105] S201, obtaining a second pedestrian image to be identified;

[0106] S202: Input the second pedestrian image into the pedestrian re-identification model obtained by the aforementioned pedestrian re-identification model training method to obtain a pedestrian re-identification result.

[0107] It can be understood that the contents of the above-mentioned pedestrian re-identification model training method embodiment are all applicable to the present pedestrian re-identification method embodiment, the functions specifically implemented by the present pedestrian re-identification method embodiment are the same as those in the above-mentioned pedestrian re-identification model training method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned pedestrian re-identification model training method embodiment.

[0108] Reference Figure 5 , an embodiment of the present invention provides a pedestrian re-identification model training system, comprising:

[0109] A global feature extraction module, used to obtain a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image;

[0110] A local feature extraction module, used to input the second global feature information into a preset mask generator to obtain a plurality of mask images, and then perform local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information;

[0111] A feature vector determination module, used to determine a local feature vector according to a plurality of first local feature information, and to determine a global feature vector according to the first global feature information;

[0112] The model training module is used to determine the training data set according to the local feature vector and the global feature vector, and input the training data set into the pre-built convolutional neural network for training to obtain a trained pedestrian re-identification model.

[0113] It can be understood that the contents of the above-mentioned pedestrian re-identification model training method embodiment are all applicable to the present pedestrian re-identification model training system embodiment, and the functions specifically implemented by the present pedestrian re-identification model training system embodiment are the same as those in the above-mentioned pedestrian re-identification model training method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned pedestrian re-identification model training method embodiment.

[0114] Reference Figure 6 , an embodiment of the present invention provides a pedestrian re-identification model training device, comprising:

[0115] at least one processor;

[0116] at least one memory for storing at least one program;

[0117] When the at least one program is executed by the at least one processor, the at least one processor implements the aforementioned pedestrian re-identification model training method.

[0118] It can be understood that the contents of the above-mentioned pedestrian re-identification model training method embodiment are all applicable to the present pedestrian re-identification model training device embodiment, and the functions specifically implemented by the present pedestrian re-identification model training device embodiment are the same as those in the above-mentioned pedestrian re-identification model training method embodiment, and the beneficial effects achieved are also the same as those achieved by the above-mentioned pedestrian re-identification model training method embodiment.

[0119] An embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned pedestrian re-identification model training method or the above-mentioned pedestrian re-identification method.

[0120] A computer-readable storage medium according to an embodiment of the present invention can execute the pedestrian re-identification model training method or the pedestrian re-identification method provided by the method embodiment of the present invention, can execute any combination of implementation steps of the method embodiment of the present invention, and has the corresponding functions and beneficial effects of the method embodiment of the present invention.

[0121] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 or Figure 4 The method shown.

[0122] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.

[0123] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0124] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0125] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0126] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.

[0127] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0128] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0129] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.

[0130] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A person re-identification model training method, characterized in that: The following steps are involved: Acquire a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image; Inputting the second global feature information into a preset mask generator to obtain a plurality of mask images, and then performing local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information; Determine a local feature vector according to the first local feature information, and determine a global feature vector according to the first global feature information; Determining a training data set according to the local feature vector and the global feature vector, and inputting the training data set into a pre-built convolutional neural network for training to obtain a trained person re-identification model; The step of determining the first global feature information and the second global feature information of the first pedestrian image is specifically as follows: Extracting first global feature information and second global feature information of the first pedestrian image through a global feature extraction network; The width of the first global feature information is the same as the width of the second global feature information, and the height of the first global feature information is the same as the height of the second global feature information; The step of determining a local feature vector according to the plurality of first local feature information and determining a global feature vector according to the first global feature information specifically includes: Performing global weighted pooling on the plurality of first local feature information respectively to obtain a plurality of second local feature information; Determine a plurality of first feature vectors according to a plurality of the second local feature information, and then perform splicing processing on the plurality of the first feature vectors to obtain a local feature vector; Perform global maximum pooling on the first global feature information to obtain third global feature information, and then determine a global feature vector according to the third global feature information.

2. The method for training a person re-identification model according to claim 1, characterized in that: The step of extracting local features from the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information is specifically as follows: A tensor product operation is performed on the mask image and the first global feature information to obtain a plurality of first local feature information.

3. The method for training a person re-identification model according to claim 1, characterized in that: The step of determining a training data set according to the local feature vector and the global feature vector specifically includes: Determine a training sample according to the local feature vector and the global feature vector; Acquire annotation information of the first pedestrian image, and generate a classification label according to the annotation information; A training data set is determined according to the training samples and corresponding classification labels.

4. The method for training a person re-identification model according to claim 3, characterized in that: The step of inputting the training data set into a pre-built convolutional neural network for training to obtain a trained person re-identification model specifically includes: Inputting the training data set into the convolutional neural network to obtain a prediction classification result; Determining a training loss value according to the predicted classification result and the classification label; Updating the parameters of the convolutional neural network layer by layer through a back propagation algorithm according to the loss value; When the loss value reaches a preset first threshold or the number of iterations reaches a preset second threshold, the training is stopped to obtain the pedestrian re-identification model.

5. A pedestrian re-identification method, characterized in that: The following steps are involved: Obtain a second person image to be identified; The second pedestrian image is input into the pedestrian re-identification model obtained by the pedestrian re-identification model training method according to any one of claims 1 to 4 to obtain a pedestrian re-identification result.

6. A pedestrian re-identification model training system, characterized in that: include: A global feature extraction module, used to obtain a preset first pedestrian image, and determine first global feature information and second global feature information of the first pedestrian image; A local feature extraction module, used for inputting the second global feature information into a preset mask generator to obtain a plurality of mask images, and then performing local feature extraction on the first global feature information according to the plurality of mask images to obtain a plurality of first local feature information; a feature vector determination module, configured to determine a local feature vector according to the plurality of first local feature information, and determine a global feature vector according to the first global feature information; A model training module, used to determine a training data set according to the local feature vector and the global feature vector, and input the training data set into a pre-built convolutional neural network for training to obtain a trained person re-identification model; The determining of the first global feature information and the second global feature information of the first pedestrian image is specifically: Extracting first global feature information and second global feature information of the first pedestrian image through a global feature extraction network; The width of the first global feature information is the same as the width of the second global feature information, and the height of the first global feature information is the same as the height of the second global feature information; The determining of a local feature vector according to the plurality of first local feature information, and determining a global feature vector according to the first global feature information specifically includes: Performing global weighted pooling on the plurality of first local feature information respectively to obtain a plurality of second local feature information; Determine a plurality of first feature vectors according to a plurality of the second local feature information, and then perform splicing processing on the plurality of the first feature vectors to obtain a local feature vector; Perform global maximum pooling on the first global feature information to obtain third global feature information, and then determine a global feature vector according to the third global feature information.

7. A pedestrian re-identification model training device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a pedestrian re-identification model training method as described in any one of claims 1 to 4.

8. A computer-readable storage medium storing a program executable by a processor, characterized in that: The processor-executable program is used to perform the method according to any one of claims 1 to 5 when executed by the processor.

Citation Information

Patent Citations

  • Pedestrian re-identification model training method and device and pedestrian re-identification method and device

    CN111738090A