A method and system for pedestrian re-identification

By using the backbone network to extract features in the pedestrian recognition system, and combining advanced pooling and feature fusion technology, the similarity weight is dynamically adjusted, which solves the problem of pedestrian pose non-alignment, and improves the identity matching performance and generalization ability of the model.

CN118015652BActive Publication Date: 2025-05-23SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311521110.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-23
Estimated Expiration
2043-11-13

AI Technical Summary

Technical Problem

When the existing pedestrian re-identification method deals with the problem of pedestrian posture non-alignment, it is difficult to adapt to real and complex video surveillance scenarios, and it depends on component annotation information, making it difficult to obtain sufficient annotation data in the real environment.

Method used

By preprocessing the image data, the backbone network features are extracted, and the features are mapped into multiple projection features through a convolutional layer with non-sharing parameters. Finally, higher-order pooling and feature fusion are performed, and the similarity weight is dynamically adjusted to improve the recognition performance of the aligned components.

Benefits of technology

It improves the identity matching performance and generalization capabilities of the pedestrian re-identification model, reduces the dependence on component annotation information, and can more effectively deal with pedestrian pose non-alignment problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015652B_ABST
    Figure CN118015652B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a method and system for pedestrian re-identification, the method comprising: obtaining target image data by pre-processing image data; inputting the target image data into the backbone network of the pedestrian re-identification model for feature extraction to obtain a first feature; performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features; performing high-order pooling processing on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-indexes; obtaining target high-order features by performing feature fusion on the first high-order features and the second high-order features; obtaining calculation results by calculating the similarity between the target high-order features and each candidate high-order feature; and determining the target pedestrian to which the pedestrian in the target image data belongs according to the calculation results. The purpose is to improve the identity matching performance and generalization ability of the pedestrian re-identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and in particular to a method and system for pedestrian re-identification. Background Art

[0002] Person re-identification is a cutting-edge direction and research hotspot in the field of computer vision. It aims to use computer vision technology to retrieve specific pedestrians in image or video libraries, thus providing a basis for many applications such as cross-camera target tracking and cross-camera behavior analysis. Therefore, pedestrian re-identification research has a strong practical demand and has important application value in a large number of fields.

[0003] In video surveillance systems, pedestrian targets are generally active under the camera. Collecting pedestrian samples under multiple different cameras results in large intra-class differences among pedestrians, that is, the same pedestrians have very different appearances, and only some local areas have the same details. In addition, different pedestrians wear the same or similar clothes or accessories, so the appearance differences between different pedestrians become very small, which poses a great obstacle to distinguishing the identities of different pedestrians. Pedestrian targets are non-rigid. Unlike some rigid targets, they enter the surveillance system from different directions, which causes a variety of posture misalignment problems in pedestrian images (such as changes in shape and position of various parts of the body, posture changes, background changes, human occlusion, detection errors, etc., and at the same time, the appearance of pedestrians is also affected by factors such as camera viewing angle, lighting conditions, body occlusion, and camera parameter differences), which makes it difficult to represent pedestrian features, thereby affecting the performance of pedestrian re-identification. Considering that local features have a certain robustness to intra-class changes, many researchers usually divide pedestrian images into different sub-regions and then extract local features of each sub-region when describing the appearance features of pedestrians. Therefore, to address the problem of pedestrian posture misalignment, current person re-identification methods can be divided into two types: component prior information learning and component supervised information learning.

[0004] For the first method mentioned above, it only uses pedestrian category labels to train the network, and does not use any additional pedestrian component annotation information. It divides the human body area based on the prior information of the component spatial distribution. This method can only achieve coarse-grained component alignment, but cannot perform fine semantic alignment of local features. The real complex environment where pedestrians are located often causes the pedestrian images captured by the camera to have rich posture changes, so this method may be difficult to adapt to real and complex video surveillance scenes, which will seriously affect the retrieval performance of the pedestrian re-identification model. The second method requires the use of additional pedestrian component supervision information to assist the pedestrian re-identification model in learning pedestrian component features with posture invariance. This method relies heavily on component annotation information, but it is difficult to obtain sufficient pedestrian images with component annotations and high-precision component detection networks in real environments. Low-precision component annotation and component detection will introduce additional errors, affecting the quality of the extracted local features, and may not be well generalized to pedestrian images with new changes. Summary of the invention

[0005] In view of this, an embodiment of the present invention provides a method and system for pedestrian re-identification, aiming to solve the problem of pedestrian posture misalignment so as to improve the identity matching performance and generalization ability of the pedestrian re-identification model.

[0006] A first aspect of an embodiment of the present invention provides a method for pedestrian re-identification, the method comprising:

[0007] By preprocessing the image data, target image data is obtained;

[0008] Inputting the target image data into the backbone network of the pedestrian re-identification model for feature extraction to obtain a first feature;

[0009] Performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features;

[0010] By performing high-order pooling processing on the preset number of second features, first high-order features with the same index and second high-order features with cross-indexes are obtained;

[0011] Obtaining a target high-order feature by performing feature fusion on the first high-order feature and the second high-order feature;

[0012] The target high-order feature is calculated by calculating the similarity with each candidate high-order feature to obtain a calculation result;

[0013] The target pedestrian to which the pedestrian in the target image data belongs is determined according to the calculation result.

[0014] Optionally, obtaining target image data by preprocessing the image data includes:

[0015] Processing the image data into a target size to obtain first image data;

[0016] The target image data is obtained by performing normalization processing on the first image data.

[0017] Optionally, performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features includes:

[0018] Inputting the first features into a preset number of convolutional layers for processing respectively to obtain a preset number of first sub-features;

[0019] A preset number of second features are obtained by performing batch normalization processing on the preset number of first sub-features respectively.

[0020] Optionally, performing high-order pooling processing on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-indexes includes:

[0021] Calculating the preset number of second features by a first high-order algorithm to obtain first high-order features with the same index;

[0022] Obtaining a preset number of third features by performing channel shuffling processing on the preset number of second features;

[0023] Calculating the preset number of third features by a second high-order algorithm to obtain cross-indexed second high-order features;

[0024] Wherein, the first high-order algorithm is:

[0025]

[0026] Among them, x s is the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number;

[0027] Wherein, the second high-order algorithm is:

[0028]

[0029] Among them, x c is the second highest-order feature of the cross-index, ⊙ represents the Hadamard product, Represents the i-th second convolution feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number.

[0030] Optionally, obtaining a preset number of third features by performing channel shuffling processing on the preset number of second features includes:

[0031] By performing convolution processing on the preset number of second features respectively, query features and key features corresponding to each second feature are obtained;

[0032] Obtaining an interactive attention feature corresponding to the same second feature by performing inner product processing on a query feature and a corresponding key feature corresponding to the same second feature;

[0033] By performing matrix multiplication on the interactive attention feature and the same second feature, a third feature corresponding to the same second feature is obtained.

[0034] Optionally, calculating the preset number of third features by using a second high-order algorithm to obtain cross-indexed second high-order features includes:

[0035] Calculating the preset number of third features a preset number of times by a second high-order algorithm to obtain a corresponding preset number of cross-indexed second high-order features;

[0036] The preset number of cross-indexed second high-order features are fused through a preset algorithm to obtain final cross-indexed second high-order features.

[0037] Optionally, the training of the person re-identification model includes:

[0038] Performing random size processing on the training sample image to obtain a first training sample image;

[0039] Processing the first training sample image into a target size to obtain a second training sample image;

[0040] Obtaining target training sample data by normalizing the second training sample image;

[0041] Construct an initial person re-identification model, set model parameters and corresponding update algorithms, and obtain the first person re-identification model;

[0042] Training the first person re-identification model using the target training sample data, optimizing model parameters using a joint loss function, and determining whether the initial person re-identification model is trained properly;

[0043] If the training is qualified, a second person re-identification model is obtained, and the second person re-identification model is tested by using a test sample image to obtain a test result;

[0044] When the test result indicates that the recognition of the test sample image is correct, the second person re-recognition model is determined as a finally applicable person re-recognition model.

[0045] Optionally, set model parameters and corresponding update algorithms, including:

[0046] Initializing the network weight parameters of the initial person re-identification model through the ImageNet pre-training model;

[0047] The network weight parameters are updated during the training process using the stochastic gradient descent algorithm.

[0048] Optionally, the preset joint loss function includes: a triplet loss function and an attention regularization function.

[0049] A second aspect of the present invention provides a system for pedestrian re-identification, the system comprising:

[0050] A preprocessing module, used for obtaining target image data by preprocessing the image data;

[0051] A feature extraction module, used for inputting the target image data into a backbone network of a person re-identification model for feature extraction to obtain a first feature;

[0052] A convolution processing module, used for performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features;

[0053] A high-order pooling module, used to obtain first high-order features with the same index and second high-order features with cross-indexes by performing high-order pooling processing on the preset number of second features;

[0054] A feature fusion module, used for obtaining a target high-order feature by fusing the first high-order feature and the second high-order feature;

[0055] A similarity calculation module, used for calculating the similarity between the target high-order feature and each candidate high-order feature to obtain a calculation result;

[0056] The pedestrian re-identification module is used to determine the target pedestrian to which the pedestrian in the target image data belongs based on the calculation result.

[0057] The method for pedestrian re-identification provided by the present invention has the following advantages:

[0058] A method for pedestrian re-identification provided by an embodiment of the present invention first obtains target image data by preprocessing image data; inputs the target image data into the backbone network of the pedestrian re-identification model for feature extraction to obtain a first feature; performs convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features; performs high-order pooling processing on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-indexes; obtains target high-order features by performing feature fusion on the first high-order features and the second high-order features; obtains calculation results by calculating the similarity between the target high-order features and each candidate high-order feature; and determines the target pedestrian to which the pedestrian in the target image data belongs according to the calculation results. The present invention first extracts pedestrian features (i.e., the first features) through a backbone network, then maps the pedestrian features into multiple different pedestrian projection features (i.e., the second features) through a convolutional layer with non-shared parameters, and finally inputs the multiple pedestrian projection features into a high-order compression pooling layer for processing to output high-order pedestrian features, thereby improving the identity matching performance and generalization performance of the pedestrian re-identification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative labor.

[0060] Figure 1 A flowchart of a method for pedestrian re-identification is shown in accordance with an embodiment of the present invention;

[0061] Figure 2 A schematic diagram of calculating the first high-order features with the same index in a method for pedestrian re-identification according to an embodiment of the present invention;

[0062] Figure 3 A schematic diagram of calculating a cross-indexed second high-order feature in a method for pedestrian re-identification according to an embodiment of the present invention;

[0063] Figure 4 A framework diagram of a pedestrian re-identification model in a pedestrian re-identification method according to an embodiment of the present invention;

[0064] Figure 5 The figure is a schematic diagram of a pedestrian re-identification system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] Before describing the present invention, a background description is given. Pedestrian re-identification is a technology that uses computer vision methods to retrieve the identity of a specific pedestrian in an image or video library. It aims to identify a target pedestrian that appears in a video image of a camera and then reappears in cameras at other times and different locations. Specifically, when the pedestrian target that appears in camera A disappears from the current video, the task of pedestrian re-identification is to search for the pedestrian target in camera B that does not overlap with camera A and find the pedestrian target again.

[0067] Person re-identification is a cutting-edge direction and research hotspot in the field of computer vision. It aims to use computer vision technology to retrieve specific pedestrians in image or video libraries, thus providing a basis for many applications such as cross-camera target tracking and cross-camera behavior analysis. Therefore, pedestrian re-identification research has a strong practical demand and has important application value in a large number of fields.

[0068] In video surveillance systems, pedestrian targets are generally active under the camera. Collecting pedestrian samples under multiple different cameras results in large intra-class differences among pedestrians, that is, the same pedestrians have very different appearances, and only some local areas have the same details. In addition, different pedestrians wear the same or similar clothes or accessories, so the appearance differences between different pedestrians become very small, which poses a great obstacle to distinguishing the identities of different pedestrians. Pedestrian targets are non-rigid. Unlike some rigid targets, they enter the surveillance system from different directions, which causes a variety of posture misalignment problems in pedestrian images (such as changes in shape and position of various parts of the body, posture changes, background changes, human occlusion, detection errors, etc., and at the same time, the appearance of pedestrians is also affected by factors such as camera viewing angle, lighting conditions, body occlusion, and camera parameter differences), which makes it difficult to represent pedestrian features, thereby affecting the performance of pedestrian re-identification. Considering that local features have a certain robustness to intra-class changes, many researchers usually divide pedestrian images into different sub-regions and then extract local features of each sub-region when describing the appearance features of pedestrians. Therefore, to address the problem of pedestrian posture misalignment, current person re-identification methods can be divided into two types: component prior information learning and component supervised information learning.

[0069] For the first method mentioned above, it only uses pedestrian category labels to train the network, and does not use any additional pedestrian component annotation information. It divides the human body area based on the prior information of the component spatial distribution. This method can only achieve coarse-grained component alignment, but cannot perform fine semantic alignment of local features. The real complex environment where pedestrians are located often causes the pedestrian images captured by the camera to have rich posture changes, so this method may be difficult to adapt to real and complex video surveillance scenes, which will seriously affect the retrieval performance of the pedestrian re-identification model. The second method requires the use of additional pedestrian component supervision information to assist the pedestrian re-identification model in learning pedestrian component features with posture invariance. This method relies heavily on component annotation information, but it is difficult to obtain sufficient pedestrian images with component annotations and high-precision component detection networks in real environments. Low-precision component annotation and component detection will introduce additional errors, affecting the quality of the extracted local features, and may not be well generalized to pedestrian images with new changes. In view of this, the present invention provides a method for pedestrian re-identification, which mainly solves the problem of pedestrian posture non-alignment. Without relying on pedestrian images with human body part annotations and a high-precision part detection network, the similarity weights of the aligned part pairs are dynamically improved by utilizing the similarity difference between the aligned part pairs and the non-aligned part pairs, so as to learn the pedestrian features with posture alignment and enhance the retrieval performance and generalization ability of the pedestrian re-identification model.

[0070] refer to Figure 1 , Figure 1 The flowchart of a method for pedestrian re-identification is shown in an embodiment of the present invention. The method for pedestrian re-identification provided by the present invention is as follows: Figure 1 As shown, the method includes:

[0071] Step S1: Obtain target image data by preprocessing the image data.

[0072] In this embodiment, the image data collected by the camera is first preprocessed to be processed into target image data of a target size, wherein the target size is preferably 384×128, and it should be understood that this is only a preferred implementation of the target size, and the target size can also be set to other sizes.

[0073] Step S2: input the target image data into the backbone network of the pedestrian re-identification model for feature extraction to obtain the first feature.

[0074] In this embodiment, after obtaining the target image data by preprocessing the collected image data in step S1, the target image data is input into a pre-trained pedestrian re-identification model, and the input target image data is subjected to feature extraction through the backbone network in the pedestrian re-identification model to obtain the pedestrian convolution feature in the target image data, that is, the first feature. Among them, the backbone network in the pedestrian re-identification model of the present invention is preferably a deep residual network ResNet50. It should be understood that this is only a preferred implementation of the backbone network in the pedestrian re-identification model of the present invention. The backbone network in the pedestrian re-identification model of the present invention can also be other backbone networks. At the same time, in order to increase the spatial size of the convolution feature, the present invention adjusts the step size of the last downsampling convolution layer in the deep residual network ResNet50 from 2 to 1.

[0075] Step S3: performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features.

[0076] In this embodiment, in step S2, feature extraction is performed on the input target image data through the backbone network in the pedestrian re-identification model. After obtaining the first feature in the target image data, the first feature is convolved respectively through a preset number of convolutional layers whose parameters are not shared in the pedestrian re-identification model to project the first feature to a preset number of different semantic spaces, so as to obtain a preset number of different pedestrian projection features, that is, a preset number of second features. Among them, the preset number is preferably three, and it should be understood that this is only a preferred implementation of the preset number, and the preset number can also be set to other numbers. The present invention extracts different levels of pedestrian posture information from the pedestrian convolutional features (that is, the first features) by utilizing multiple different convolutional layers whose parameters are not shared, which facilitates the high-order semantic alignment of subsequent pedestrian components.

[0077] Step S4: performing high-order pooling processing on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-index.

[0078] In this embodiment, after obtaining a preset number of second features in step S3, high-order pooling is performed on the preset number of second features to obtain a first high-order feature with the same index and a second high-order feature with a cross index.

[0079] Step S5: Obtain target high-order features by fusing the first high-order features and the second high-order features.

[0080] In this embodiment, after obtaining a first high-order feature with the same index and a second high-order feature with a cross-index through high-order pooling processing in step S4, a feature concatenation operation is performed on the first high-order feature and the second high-order feature to fuse the first high-order feature and the second high-order feature, thereby enriching the identification semantic information of the high-order feature, thereby obtaining the target high-order feature after the feature fusion of the first high-order feature and the second high-order feature.

[0081] Step S6: Calculate the similarity between the target high-order feature and each candidate high-order feature to obtain a calculation result.

[0082] In this embodiment, after obtaining the target high-order feature in step S5, the target high-order feature is respectively calculated for similarity with each candidate high-order feature to obtain a calculation result. A candidate high-order feature and a target high-order feature are calculated for similarity, and there are as many corresponding calculation results as the number of candidate high-order features. For example, if there are 500 candidate high-order features, there are 500 corresponding calculation results, and one calculation result corresponds to one candidate high-order feature, and the one calculation result represents the similarity between the target high-order feature and the candidate high-order feature. Among them, the implementation method of processing the image data corresponding to the candidate high-order feature into the candidate high-order feature is the same as the implementation method of processing the image data corresponding to the above-mentioned target high-order feature into the target high-order feature; at the same time, the camera that collects the image data corresponding to the candidate high-order feature can be a camera different from the camera that collects the image data corresponding to the target high-order feature, or it can be the same camera as the camera that collects the image data corresponding to the target high-order feature.

[0083] Step S7: Determine the target pedestrian to which the pedestrian in the target image data belongs based on the calculation result.

[0084] In this embodiment, various similarity calculation results obtained by comparison calculation are determined to determine the calculation result with the highest similarity, and then it is determined which candidate high-order feature corresponds to the calculation result with the highest similarity, and the pedestrian in the image data corresponding to the candidate high-order feature corresponding to the calculation result with the highest similarity is determined as the target pedestrian. At this time, the pedestrian in the image data corresponding to the target high-order feature is determined to be the target pedestrian, that is, it is determined that the pedestrian in the image data corresponding to the target high-order feature and the pedestrian in the image data corresponding to the candidate high-order feature corresponding to the calculation result with the highest similarity belong to the same person.

[0085] A method for pedestrian re-identification provided by an embodiment of the present invention first obtains target image data by preprocessing image data; inputs the target image data into the backbone network of the pedestrian re-identification model for feature extraction to obtain a first feature; performs convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features; performs high-order pooling processing on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-indexes; obtains target high-order features by performing feature fusion on the first high-order features and the second high-order features; obtains calculation results by calculating the similarity between the target high-order features and each candidate high-order feature; and determines the target pedestrian to which the pedestrian in the target image data belongs according to the calculation results. The present invention first extracts pedestrian features (i.e., the first features) through a backbone network, then maps the pedestrian features into multiple different pedestrian projection features (i.e., the second features) through a convolutional layer with non-shared parameters, and finally inputs the multiple pedestrian projection features into a high-order compression pooling layer for processing to output high-order pedestrian features, thereby improving the identity matching performance and generalization performance of the pedestrian re-identification model. Here, the present invention increases the numerical difference between the similarity of the aligned parts and the similarity of the non-aligned parts of the pedestrians by utilizing a high-order mapping function, thereby highlighting the importance of the similarity of the aligned parts, which helps the model learn the characteristics of pedestrians with aligned postures.

[0086] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, step S1 includes steps S11 to S12:

[0087] Step S11: Process the image data into a target size to obtain first image data.

[0088] In this embodiment, the image data collected by the camera is first processed into first image data of a target size, wherein the target size is preferably 384×128, and it should be understood that this is only a preferred implementation of the target size, and the target size can also be set to other sizes.

[0089] Step S12: obtaining target image data by normalizing the first image data.

[0090] In this embodiment, after the image data collected by the camera is processed into the first image data of the target size through step S11, the obtained first image data is normalized to obtain the corresponding target image data. Among them, a preferred implementation of the normalization process is to perform the normalization process by the normalization method of the ImageNet pre-training model. Specifically, for the input first image data, the mean [0.485, 0.456, 0.406] is first subtracted pixel by pixel, and then divided by the standard deviation value [0.229, 0.224, 0.225] pixel by pixel to perform an image normalization operation on the input first image to obtain the corresponding target image data. It should be understood that this is only a preferred implementation of the normalization process, and the normalization process can also be other normalization processing operations, which are not specifically limited here.

[0091] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, step S3 includes steps S31 to S32:

[0092] Step S31: input the first features into a preset number of convolutional layers for processing to obtain a preset number of first sub-features.

[0093] In this embodiment, the first feature obtained in step S2 is first convolved through a preset number of convolutional layers whose parameters are not shared in the pedestrian re-identification model to obtain a corresponding preset number of first sub-features.

[0094] Step S32: Obtain a preset number of second features by performing batch normalization processing on the preset number of first sub-features respectively.

[0095] In this embodiment, after obtaining a preset number of first sub-features through convolution processing in step S31, batch normalization processing is performed on the preset number of first sub-features to obtain a corresponding preset number of second features.

[0096] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, step S4 includes steps S41 to S43:

[0097] Step S41: Calculate the preset number of second features using a first high-order algorithm to obtain first high-order features with the same index;

[0098] Wherein, the first high-order algorithm is:

[0099]

[0100] Among them, x sis the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number.

[0101] In this embodiment, the first high-order features with the same index are obtained by calculating the preset number of second features obtained by the first high-order algorithm.

[0102] In this embodiment, the first high-order algorithm is:

[0103]

[0104] Among them, x s is the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V x represents the location set of all component features, |•| represents the set capacity, and n is the preset number.

[0105] For example, Figure 2 As shown in the figure, positions 1 to 4 refer to four parts of the target image data, that is, the target image data is divided into four parts according to the coordinate position of the target image data itself, and positions 1 to 4 together constitute the target image data (wherein, the number of parts into which the target image data is divided can be any number, and 4 is used as an example here). If the preset number is set to 2, the number of second features is two, which are represented by second feature A and second feature B. Figure 2 a1 to a4 in constitute the second feature A of the two second features, such as Figure 2 b1 to b4 in the above constitute another second feature B of the two second features. The second feature A and the second feature B belong to the features of different channels of the same target image data. By performing the above steps S2 and S3 on the target image data, the two second features, namely the second feature A and the second feature B, are obtained. Figure 2Fa1 represents the component feature of the second feature A at position 1, Fa2 represents the component feature of the second feature A at position 2, Fa3 represents the component feature of the second feature A at position 3, and Fa4 represents the component feature of the second feature A at position 4. These four component features together constitute the second feature A. Fb1 represents the component feature of the second feature B at position 1, Fb2 represents the component feature of the second feature B at position 2, Fb3 represents the component feature of the second feature B at position 3, and Fb4 represents the component feature of the second feature B at position 4. These four component features together constitute the second feature B.

[0106] Since the preset number of second features based on the first high-order features of the same index are not subjected to channel shuffle processing, the component feature Fa1 of the second feature A at position 1 is converted from the image data of the corresponding target image data at position 1, the component feature Fa2 of the second feature A at position 2 is converted from the image data of the corresponding target image data at position 2, the component feature Fa3 of the second feature A at position 3 is converted from the image data of the corresponding target image data at position 3, and the component feature Fa4 of the second feature A at position 4 is converted from the image data of the corresponding target image data at position 4. At the same time, the component feature Fb1 of the second feature b at position 1 is converted from the image data of the corresponding target image data at position 1, the component feature Fb2 of the second feature b at position 2 is converted from the image data of the corresponding target image data at position 2, the component feature Fb3 of the second feature b at position 3 is converted from the image data of the corresponding target image data at position 3, and the component feature Fb4 of the second feature b at position 4 is converted from the image data of the corresponding target image data at position 4. Therefore, the specific calculation process of calculating the second feature A and the second feature B by the first high-order algorithm is: Figure 2 As shown, the Hadamard product is performed on the component feature Fa1 of the second feature A at position 1 and the component feature Fb1 of the second feature B at position 1, the Hadamard product is performed on the component feature Fa2 of the second feature A at position 2 and the component feature Fb2 of the second feature B at position 2, the Hadamard product is performed on the component feature Fa3 of the second feature A at position 3 and the component feature Fb3 of the second feature B at position 3, the Hadamard product is performed on the component feature Fa4 of the second feature A at position 4 and the component feature Fb4 of the second feature B at position 4, and the four Hadamard products are added to obtain the corresponding first high-order features with the same index.

[0107] In this embodiment, since the dimension of the high-order features with the same index grows linearly, if the target image data is divided into 100 parts, the corresponding same index amount is only 100, that is, each second feature is indexed by the component features at its respective position 1. Therefore, when calculating the first high-order features with the same index, the present invention does not compress the first high-order features with the same index, but directly calculates a preset number of second features to obtain the first high-order features with the same index.

[0108] Step S42: Obtain a preset number of third features by performing channel shuffling processing on the preset number of second features.

[0109] In this embodiment, before determining the corresponding cross-indexed second high-order features based on the obtained preset number of second features, channel shuffling processing is first performed on the preset number of second features to obtain a preset number of third features.

[0110] For example, Figure 3 As shown in the figure, positions 1 to 4 refer to four parts of the target image data, that is, the target image data is divided into four parts according to the coordinate position of the target image data itself, and positions 1 to 4 together constitute the target image data (wherein, the number of parts into which the target image data is divided can be any number, and 4 is used as an example here). If the preset number is set to 2, the number of second features is two, which are represented by second feature A and second feature B. Figure 2 a1 to a4 in constitute the second feature A of the two second features, such as Figure 2 b1 to b4 in the above constitute another second feature B of the two second features. The second feature A and the second feature B belong to the features of different channels of the same target image data. By performing the above steps S2 and S3 on the target image data, the two second features, namely the second feature A and the second feature B, are obtained. Figure 2 Fa1 represents the component feature of the second feature A at position 1, Fa2 represents the component feature of the second feature A at position 2, Fa3 represents the component feature of the second feature A at position 3, and Fa4 represents the component feature of the second feature A at position 4. These four component features together constitute the second feature A. Fb1 represents the component feature of the second feature B at position 1, Fb2 represents the component feature of the second feature B at position 2, Fb3 represents the component feature of the second feature B at position 3, and Fb4 represents the component feature of the second feature B at position 4. These four component features together constitute the second feature B.

[0111] Channel shuffling processing is performed on a preset number of second features. In this example, channel shuffling processing is performed on each of the second features A and B, and finally the channel shuffling processing result for the second feature A is obtained, and the channel shuffling processing result for the second feature B is obtained. For the channel shuffling processing result for the second feature A, Figure 3 As shown, the component feature F′a3 of the third feature A′ at position 1 is obtained, and most of the component feature F′a3 is composed of the component feature Fa3 of the second feature A at position 3 before the channel shuffle processing is performed, and the component feature Fa3 of the second feature A at position 3 before the channel shuffle processing is performed is the image data converted from the corresponding target image data at position 3; the component feature F′a2 of the third feature A′ at position 2 is obtained, and most of the component feature F′a2 is composed of the component feature Fa2 of the second feature A at position 2 before the channel shuffle processing is performed, and the component feature Fa2 of the second feature A at position 2 before the channel shuffle processing is the image data converted from the corresponding target image data at position 2 ; the component feature F′a4 at position 3 of the third feature A′ is obtained, most of which is composed of the component feature Fa4 at position 4 of the second feature A before the channel shuffle processing is performed, and the component feature Fa4 at position 4 of the second feature A before the channel shuffle processing is converted from the image data at position 4 of the corresponding target image data; the component feature F′a1 at position 4 of the third feature A′ is obtained, most of which is composed of the component feature Fa1 at position 1 of the second feature A before the channel shuffle processing is performed, and the component feature Fa1 at position 1 of the second feature A before the channel shuffle processing is converted from the image data at position 1 of the corresponding target image data. The four component features F′a3, F′a2, F′a4, and F′a1 together constitute the third feature A′ obtained after the second feature A is processed by the channel shuffle.

[0112] The results of channel out-of-order processing for the second feature B are as follows Figure 3As shown, a component feature F′a2 of the third feature B′ at position 1 is obtained, most of which is composed of component feature Fa2 of the second feature B at position 2 before the channel shuffle processing is performed, and the rest is composed of component features Fa1, Fa3, and Fa4 of the second feature B at positions 1, 3, and 4 before the channel shuffle processing is performed, and the component feature Fa2 of the second feature B at position 2 before the channel shuffle processing is converted from the image data of the corresponding target image data at position 2; a component feature F′a4 of the third feature B′ at position 2 is obtained, most of which is composed of component feature Fa4 of the second feature B at position 4 before the channel shuffle processing is performed, and the rest is composed of component features Fa1, Fa2, and Fa3 of the second feature B at positions 1, 2, and 3 before the channel shuffle processing is performed, and the component feature Fa4 of the second feature B at position 4 before the channel shuffle processing is converted from the image data of the corresponding target image data at position 4 ; a component feature F′a1 at position 3 of the third feature B′ is obtained, most of which is composed of component feature Fa1 at position 1 of the second feature B before channel shuffling processing is performed, and the rest is composed of component features Fa2, Fa3, and Fa4 at positions 2, 3, and 4 of the second feature B before channel shuffling processing is performed, and the component feature Fa1 at position 1 of the second feature B before channel shuffling processing is converted from the image data of the corresponding target image data at position 1; a component feature F′a3 at position 4 of the third feature B′ is obtained, most of which is composed of component feature Fa3 at position 3 of the second feature B before channel shuffling processing is performed, and the rest is composed of component features Fa1, Fa2, and Fa4 at positions 1, 2, and 4 of the second feature B before channel shuffling processing is performed, and the component feature Fa3 at position 3 of the second feature B before channel shuffling processing is converted from the image data of the corresponding target image data at position 3. The four component features F′b2, F′b4, F′b1, and F′b4 together constitute the third feature B′ obtained after the second feature B is processed by channel shuffle, thereby obtaining a preset number of third features, that is, 2 third features, namely the third feature A′ and the third feature B′.

[0113] In this embodiment, since the dimension of the high-order features of the cross-index increases exponentially, if the target image data is divided into 100 parts and the preset number is 2, the corresponding cross-index amount is 100. 2, that is, 10,000. At the same time, as the number of preset numbers increases and the target image data division part increases, the corresponding cross-index amount will be larger. In some cases, the high-order features of the cross-index may not be calculated because the amount of calculation is too large. Therefore, when calculating the second high-order features of the cross-index, the present invention compresses the second high-order features of the cross-index, that is, first perform channel shuffling processing on the preset number of second features to obtain a preset number of third features, and then calculate the preset number of third features to obtain the second high-order features of the cross-index, that is, through channel shuffling processing, the purpose of compressing the second high-order features of the cross-index is achieved. For example, continuing with the above example, when the second high-order features of the cross-index are not compressed, the index amount of the cross-index is 4×4, and after the channel shuffling processing of the present invention is performed to achieve the purpose of compressing the second high-order features of the cross-index, the index amount of the cross-index is 4, such as Figure 3 Here, the high-order feature compression method of the present invention can not only compress the dimensions of the cross-indexed high-order features, but also retain the identification information of the cross-indexed high-order features as much as possible, which is helpful to improve the identification and robustness of the pedestrian re-identification model. The high-order feature compression method performs a differentiable channel disorder operation on the second feature of the input through the index attention mechanism, so as to fully retain the key cross-indexed high-order interaction relationship, which helps to optimize the algorithm space complexity from exponential level to linear level.

[0114] Step S43: Calculating the preset number of third features by a second high-order algorithm to obtain cross-indexed second high-order features;

[0115] Wherein, the second high-order algorithm is:

[0116]

[0117] Among them, x c is the second highest-order feature of the cross-index, ⊙ represents the Hadamard product, Represents the i-th second convolution feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number.

[0118] In this embodiment, after a preset number of third features are obtained in step S42, the preset number of third features obtained are calculated by a second high-order algorithm to obtain cross-indexed second high-order features.

[0119] In this embodiment, the second high-order algorithm is:

[0120]

[0121] Among them, x c is the first high-order feature with the same index, ⊙ represents the Hadamard product, Represents the i-th second convolution feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number.

[0122] For example, Figure 3 As shown, continuing with the above example, after obtaining a preset number of third features through channel shuffle processing, in this example, two third features are obtained, namely the third feature A′ and the third feature B′. The four component features F′a3, F′a2, F′a4, and F′a1 together constitute the third feature A′, and the four component features F′b2, F′b4, F′b1, and F′b4 together constitute the third feature B′. The specific calculation process of calculating the third feature A′ and the third feature B′ by the second high-order algorithm is as follows: Figure 3 As shown, the Hadamard product is performed on the component feature F′a3 of the third feature A′ at position 1 and the component feature F′b2 of the third feature B′ at position 1, the Hadamard product is performed on the component feature F′a2 of the third feature A′ at position 2 and the component feature F′b1 of the third feature B′ at position 2, the Hadamard product is performed on the component feature F′a4 of the third feature A′ at position 3 and the component feature F′b1 of the third feature B′ at position 3, the Hadamard product is performed on the component feature F′a1 of the third feature A′ at position 4 and the component feature F′b3 of the third feature B′ at position 4, and the four Hadamard products are added to obtain the corresponding cross-indexed second high-order feature.

[0123] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, step S42 includes steps S421 to S423:

[0124] Step S421: performing convolution processing on the preset number of second features respectively to obtain query features and key features corresponding to each second feature.

[0125] In this embodiment, after obtaining a preset number of second features in step S3, a preset number of different convolutional layers are used to perform convolution processing on the preset number of second features respectively to obtain query features and key features corresponding to each second feature.

[0126] For example, if the preset number is determined to be 2, two second features will be obtained, represented by the second feature A and the second feature B. At this time, the second feature A is convolved with a convolution layer x1 to obtain the query feature a1 and the key feature a2 corresponding to the second feature A. The second feature B is convolved with a convolution layer y1 to obtain the query feature b1 and the key feature b2 corresponding to the second feature B.

[0127] Step S422: Obtain the interactive attention feature corresponding to the same second feature by performing inner product processing on the query feature and the corresponding key feature corresponding to the same second feature.

[0128] In this embodiment, after obtaining a corresponding query feature and a corresponding key feature for each second feature, the query feature and the key feature corresponding to the same second feature are processed with an inner product to obtain the interactive attention feature corresponding to the same second feature. The calculation method for calculating the corresponding interactive attention feature for each second feature is the same as the calculation method for calculating the corresponding interactive attention feature for the second feature described above, and will not be repeated here.

[0129] For example, continuing with the above example, after obtaining the query feature a1 and key feature a2 corresponding to the second feature A and obtaining the query feature b1 and key feature b2 corresponding to the second feature B, perform inner product processing on the query feature a1 and the key feature a2 to obtain the interactive attention feature a1a2 corresponding to the second feature A; perform inner product processing on the query feature b1 and the key feature b2 to obtain the interactive attention feature b1b2 corresponding to the second feature B.

[0130] Step S423: Obtain a third feature corresponding to the same second feature by performing matrix multiplication on the interactive attention feature and the same second feature.

[0131] In this embodiment, after obtaining an interactive attention feature corresponding to each second feature, the interactive attention feature is matrix-multiplied with the second feature corresponding to itself, thereby obtaining the third feature corresponding to the second feature. The calculation implementation method for obtaining the corresponding third feature by calculation for each second feature is the same as the calculation implementation method for obtaining the corresponding third feature by calculation for the above-mentioned second feature, and will not be repeated here. Thus, the third feature corresponding to each second feature is obtained.

[0132] For example, continuing with the above example, after obtaining the interactive attention feature a1a2 corresponding to the second feature A and the interactive attention feature b1b2 corresponding to the second feature B, perform matrix multiplication of the interactive attention feature a1a2 and the second feature A to obtain the third feature Aa1a2 corresponding to the second feature A; perform matrix multiplication of the interactive attention feature b1b2 and the second feature B to obtain the third feature Bb1b2 corresponding to the second feature B.

[0133] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, step S43 includes steps S431 to S432:

[0134] Step S431: Calculate the preset number of third features a preset number of times using a second high-order algorithm to obtain a corresponding preset number of cross-indexed second high-order features.

[0135] In this embodiment, a preset number of calculations are performed in the same manner as the above-mentioned method of calculating a preset number of third features based on the second high-order algorithm to obtain cross-indexed second high-order features to obtain a corresponding preset number of cross-indexed second high-order features.

[0136] For example, a preset number of third features are calculated 10 times by the second high-order algorithm, thereby obtaining 10 cross-indexed second high-order features.

[0137] In this embodiment, another implementation method for obtaining a preset number of cross-indexed second high-order features is: after obtaining a preset number of second features in step S3, the preset number of second features obtained are processed respectively by a multi-head attention mechanism, thereby obtaining multiple query features and multiple key features for each second feature, and the multiple query features and the multiple key features have a one-to-one correspondence. The query features and key features with a corresponding relationship are processed by inner product to obtain the interactive attention features of a corresponding second feature. The number of query features (or key features) in a second feature can be calculated to obtain the corresponding number of interactive attention features, and each interactive attention feature is matrix-multiplied with the corresponding second feature to obtain multiple third features corresponding to the second feature. Thus, multiple corresponding third features will be calculated for each second feature, and the number of the multiple third features is the same as the number of query features in a second feature. Then, the third features obtained by the same attention mechanism of each second feature based on the multi-head attention mechanism are combined into a set, and then the third features in the same set are calculated using the second high-order algorithm to obtain a corresponding cross-indexed second high-order feature. Thus, a corresponding cross-indexed second high-order feature can be calculated for each calculation, thereby obtaining a plurality of cross-indexed second high-order features, the number of which is the same as the number of query features possessed by a second feature.

[0138] For example, when the preset number is 2, two second features are obtained, represented by the second feature A and the second feature B, and the obtained second feature A and the second feature B are processed respectively by the multi-head attention mechanism. For the second feature A, 10 query features Qa1 to Qa10 and 10 key features Ka1 to Ka10 corresponding to the 10 query features are obtained, and for the second feature B, 10 query features Qb1 to Qb10 and 10 key features Kb1 to Kb10 corresponding to the 10 query features are obtained. Qa1 and Ka1 are processed by inner product to obtain an interactive attention feature QKa1 of the second feature A, Qa2 and Ka2 are processed by inner product to obtain an interactive attention feature QKa2 of the second feature A, ..., Qa10 and Ka10 are processed by inner product to obtain an interactive attention feature QKa10 of the second feature A, thereby obtaining 10 corresponding interactive attention features QKa1 to QKa10 for the second feature A. Perform inner product processing on Qb1 and Kb1 to obtain an interactive attention feature QKb1 of the second feature B, perform inner product processing on Qb2 and Kb2 to obtain an interactive attention feature QKb2 of the second feature B, ..., perform inner product processing on Qb10 and Kb10 to obtain an interactive attention feature QKb10 of the second feature B, thereby obtaining 10 corresponding interactive attention features QKb1 to QKb10 for the second feature B. Perform matrix multiplication on the interactive attention feature QKa1 and the second feature A to obtain a corresponding third feature QKA1, perform matrix multiplication on the interactive attention feature QKa2 and the second feature A to obtain a corresponding third feature QKA2, ..., perform matrix multiplication on the interactive attention feature QKa10 and the second feature A to obtain a corresponding third feature QKA10, thereby obtaining 10 third features QKA1 to QKA10 corresponding to the second feature A. The interactive attention feature QKb1 is matrix-multiplied with the second feature B to obtain a corresponding third feature QKB1, the interactive attention feature QKb2 is matrix-multiplied with the second feature B to obtain a corresponding third feature QKB2, ..., the interactive attention feature QKb10 is matrix-multiplied with the second feature B to obtain a corresponding third feature QKB10, thereby obtaining 10 third features QKB1 to QKB10 corresponding to the second feature B. The third features QKA1 and QKB1 are calculated by the second high-order algorithm to obtain the corresponding cross-indexed second high-order feature C1, the third features QKA2 and QKB2 are calculated by the second high-order algorithm to obtain the corresponding cross-indexed second high-order feature C2, ..., the third features QKA10 and QKB10 are calculated by the second high-order algorithm to obtain the corresponding cross-indexed second high-order feature C10, thereby obtaining 10 cross-indexed second high-order features C1 to C10.Here, the present invention proposes a multi-head index attention mechanism, which uses multiple different types of index attention mechanisms to calculate multiple different cross-index second-high-order features, and performs average and maximum operations on these second-high-order features to calculate the final fused second-high-order features, which helps the pedestrian re-identification model to fully explore the visual information with discriminativeness and robustness.

[0139] Step S432: fusing the preset number of cross-indexed second high-order features using a preset algorithm to obtain final cross-indexed second high-order features.

[0140] In this embodiment, after obtaining a preset number of cross-indexed second high-order features by calculation, the preset number of cross-indexes obtained are fused through a preset algorithm to obtain the final cross-indexed second high-order features, which will be used for subsequent fusion with the first high-order features of the same index.

[0141] In this embodiment, the preset algorithm is:

[0142]

[0143] Among them, x′ c is the second highest-order feature of the final cross-index, is the second high-order feature of the j-th cross-index, the value of h is the same as the preset number, avg(·) is used to calculate the average value, and max(·) is used to calculate the maximum value.

[0144] In combination with the above embodiments, in one implementation, the present invention also provides a method for pedestrian re-identification. In the method for pedestrian re-identification, the training of the pedestrian re-identification model includes steps S01 to S08:

[0145] Step S01: Perform random size processing on a training sample image to obtain a first training sample image.

[0146] In this embodiment, the image area randomly sampled in a scale interval and an aspect ratio interval of the training sample image is cropped to obtain the first training sample image, which is to increase the randomness and diversity of the training sample. Among them, the scale interval is preferably [0.64, 1], and the aspect ratio interval is preferably [2, 3]. It should be understood that the scale interval is only a preferred implementation, and the scale interval can also be other ranges, which are not specifically limited here. At the same time, the aspect ratio interval is only a preferred implementation, and the aspect ratio interval can also be other ranges, which are not specifically limited here.

[0147] Step S02: Processing the first training sample image into a target size to obtain a second training sample image.

[0148] In this embodiment, after the training sample image is cropped to obtain the first training sample image in step S01, the size of the first training sample image is adjusted to the target size, thereby obtaining the second training sample image. At the same time, in order to increase the training data, the second training sample image can also be processed by data enhancement techniques such as horizontal flipping and random erasing to obtain more second training sample images.

[0149] Step S03: obtaining target training sample data by normalizing the second training sample image.

[0150] In this embodiment, after obtaining the second training sample image, the second training sample image is normalized to obtain the corresponding target training sample data. Among them, a preferred implementation of the normalization process is to perform the normalization process by the normalization method of the ImageNet pre-training model. Specifically, for the input first image data, the mean [0.485, 0.456, 0.406] is first subtracted pixel by pixel, and then divided pixel by pixel by the standard deviation value [0.229, 0.224, 0.225] to perform image normalization operation on the input second training sample image to obtain the corresponding target training sample data. It should be understood that this is only a preferred implementation of the normalization process, and the normalization process can also be other normalization processing operations, which are not specifically limited here. The target training sample data will be used for subsequent training of the model, and a large amount of target training sample data is used to form a model training sample set. At the same time, a large amount of other target training sample data different from the target training sample data in the training sample set is used to form a model test sample set, and the target training sample data as the test sample is not subjected to data enhancement processing.

[0151] Step S04: construct an initial person re-identification model, set model parameters and set corresponding update algorithms to obtain a first person re-identification model.

[0152] In this embodiment, an initial pedestrian re-identification model is constructed, such as Figure 4 As shown, Figure 4 The framework diagram of the pedestrian re-identification model in the present invention is shown, and the pedestrian re-identification model includes a backbone network, a convolution layer, a batch normalization layer, a high-order pooling layer, and a loss function, wherein the backbone network is preferably ResNet50. After the initial pedestrian re-identification model is constructed, the model parameters and the corresponding update algorithm are set. After the initial pedestrian re-identification model is constructed, the model parameters and the corresponding update algorithm are set, the first pedestrian re-identification model is obtained.

[0153] Step S05: training the first person re-identification model using the target training sample data, optimizing model parameters using a joint loss function, and determining whether the initial person re-identification model is qualified.

[0154] In this embodiment, after the first person re-recognition model is obtained, the first person re-recognition model is trained using a large amount of target training sample data in a model training sample set consisting of target training sample data to obtain corresponding training results.

[0155] In this embodiment, an optional implementation method of training is: within the minimum training batch, 16 pedestrian categories are randomly selected, and then 4 target training sample data are randomly selected for each pedestrian category, thereby forming a training batch containing 64 target training sample data. The model is trained for 750 epochs at the same time. After each training, the joint loss function is used to determine whether the initial pedestrian re-identification model is qualified. If it is not qualified, the model parameters are updated and trained again until the joint loss function is used to determine whether the initial pedestrian re-identification model is qualified.

[0156] Step S07: if the training is qualified, a second person re-recognition model is obtained, and the second person re-recognition model is tested through a test sample image to obtain a test result.

[0157] In this embodiment, when the initial pedestrian re-identification model is trained to be qualified, the qualified initial pedestrian re-identification model is the second pedestrian re-identification model. At this time, the second pedestrian re-identification model is tested by each test sample image in the model test sample set to obtain the test result.

[0158] Step S08: when the test result indicates that the recognition of the test sample image is correct, determining the second pedestrian re-recognition model as a finally applicable pedestrian re-recognition model.

[0159] In this embodiment, when the test results indicate that the recognition of multiple test sample images is correct, the second person re-recognition model is determined as the final applicable person re-recognition model.

[0160] In combination with the above embodiments, in one implementation, an embodiment of the present invention further provides a method for pedestrian re-identification. In the method for pedestrian re-identification, setting model parameters and setting corresponding update algorithms in step S04 include steps S041 to S042:

[0161] Step S041: Initializing the network weight parameters of the initial person re-identification model through the ImageNet pre-training model.

[0162] Step S042: Update the network weight parameters during the training process through the stochastic gradient descent algorithm.

[0163] In this embodiment, setting the model parameters and setting the corresponding update algorithm include adjusting the step size of the last downsampling convolution layer of the backbone network from 2 to 1 to increase the spatial scale of the convolution feature, initializing the network weight parameters of the model using the ImageNet pre-trained model, and updating the network weight parameters of the model using the stochastic gradient descent algorithm. At the same time, setting the experimental parameters includes: setting the initial learning rate to 0.01, setting the weight decay parameter to 2×10 -4 The momentum parameter is 0.9, and the learning rate is divided by 5 every 200 rounds of training to gradually reduce the parameter update speed.

[0164] In combination with the above embodiments, in one implementation, the present invention further provides a method for pedestrian re-identification. In the method for pedestrian re-identification, the preset joint loss function includes: a triplet loss function and an attention regularization function.

[0165] In this embodiment, the preset joint loss function in the pedestrian re-identification method provided by the present invention is composed of a triplet loss function and an attention regularization function. The triplet loss function is used to maximize the inter-class distance and minimize the intra-class distance in order to learn discriminative pedestrian features, while the attention regularization function is used to maximize the difference between different attention distributions in order to model differential cross-indexed high-order features.

[0166] Based on the same inventive concept, the second aspect of the present invention provides a pedestrian re-identification system 500, the system 500 comprising:

[0167] The preprocessing module 501 is used to obtain target image data by preprocessing the image data;

[0168] A feature extraction module 502 is used to input the target image data into the backbone network of the pedestrian re-identification model to extract features and obtain a first feature;

[0169] A convolution processing module 503, configured to perform convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features;

[0170] A high-order pooling module 504 is used to obtain first high-order features with the same index and second high-order features with cross-indexes by performing high-order pooling processing on the preset number of second features;

[0171] A feature fusion module 505, configured to obtain a target high-order feature by fusing the first high-order feature and the second high-order feature;

[0172] A similarity calculation module 506 is used to obtain a calculation result by calculating the similarity between the target high-order feature and each candidate high-order feature respectively;

[0173] The pedestrian re-identification module 507 is used to determine the target pedestrian to which the pedestrian in the target image data belongs according to the calculation result.

[0174] Optionally, the preprocessing module 501 includes:

[0175] A first preprocessing module, used for processing the image data into a target size to obtain first image data;

[0176] The normalization module is used to obtain target image data by performing normalization processing on the first image data.

[0177] Optionally, the convolution processing module 503 includes:

[0178] A convolution processing submodule, used for inputting the first features into a preset number of convolution layers for processing respectively, to obtain a preset number of first sub-features;

[0179] The batch normalization module is used to obtain a preset number of second features by performing batch normalization processing on the preset number of first sub-features respectively.

[0180] Optionally, the high-order pooling module 504 includes:

[0181] A first high-order pooling module, used to calculate the preset number of second features by using a first high-order algorithm to obtain first high-order features with the same index;

[0182] a channel disorder processing module, used for obtaining a preset number of third features by performing channel disorder processing on the preset number of second features;

[0183] A second high-order pooling module, used to calculate the preset number of third features by a second high-order algorithm to obtain cross-indexed second high-order features;

[0184] Wherein, the first high-order algorithm is:

[0185]

[0186] Among them, x s is the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V xrepresents the location set of all component features, |·| represents the set capacity, and n is the preset number;

[0187] Wherein, the second high-order algorithm is:

[0188]

[0189] Among them, x c is the second highest-order feature of the cross-index, ⊙ represents the Hadamard product, Represents the i-th second convolution feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number.

[0190] Optionally, the channel disorder processing module includes:

[0191] A first convolution module is used to obtain query features and key features corresponding to each second feature by performing convolution processing on the preset number of second features respectively;

[0192] An inner product module, used for obtaining an interactive attention feature corresponding to the same second feature by performing inner product processing on a query feature and a corresponding key feature corresponding to the same second feature;

[0193] A matrix multiplication module is used to obtain a third feature corresponding to the same second feature by performing matrix multiplication on the interactive attention feature and the same second feature.

[0194] Optionally, the second high-order pooling module includes:

[0195] A second high-order pooling submodule is used to calculate the preset number of third features a preset number of times by using a second high-order algorithm to obtain a corresponding preset number of cross-indexed second high-order features;

[0196] The feature fusion module is used to fuse the preset number of cross-indexed second high-order features through a preset algorithm to obtain the final cross-indexed second high-order features.

[0197] Optionally, the training of the person re-identification model in the feature extraction module 502 includes:

[0198] A random size processing module, used for performing random size processing on the training sample image to obtain a first training sample image;

[0199] A target size processing module, used for processing the first training sample image into a target size to obtain a second training sample image;

[0200] A first normalization processing module, used for obtaining target training sample data by performing normalization processing on the second training sample image;

[0201] A model building module is used to build an initial person re-identification model, set model parameters and set a corresponding update algorithm to obtain a first person re-identification model;

[0202] A model training module, used to train the first person re-recognition model using the target training sample data, optimize model parameters using a joint loss function, and determine whether the initial person re-recognition model is trained properly;

[0203] A model testing module, used for obtaining a second person re-identification model if the training is qualified, and testing the second person re-identification model through a test sample image to obtain a test result;

[0204] The model determination module can also determine the second pedestrian re-recognition model as the final applicable pedestrian re-recognition model when the test result indicates that the recognition of the test sample image is correct.

[0205] Optionally, the model building module includes:

[0206] A parameter initialization module, used to initialize the network weight parameters of the initial person re-identification model through the ImageNet pre-training model;

[0207] The parameter update module is used to update the network weight parameters during the training process through the stochastic gradient descent algorithm.

[0208] Optionally, the preset joint loss function in the model training module includes: a triplet loss function and an attention regularization function.

[0209] As for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0210] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0211] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0212] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, embodiments of the present invention may take the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware. Moreover, embodiments of the present invention may take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0213] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0214] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0215] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0216] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0217] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.

[0218] The above has introduced in detail a method and system for pedestrian re-identification provided by the present invention. Specific examples are used in this text to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for pedestrian re-identification, It is characterized in that The method comprises: By preprocessing the image data, target image data is obtained; Inputting the target image data into a backbone network of a person re-identification model for feature extraction to obtain a first feature; Performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features; By performing high-order pooling processing on the preset number of second features, first high-order features with the same index and second high-order features with cross-indexes are obtained; Obtaining a target high-order feature by performing feature fusion on the first high-order feature and the second high-order feature; The target high-order feature is calculated by calculating the similarity with each candidate high-order feature to obtain a calculation result; Determining, according to the calculation result, the target pedestrian to which the pedestrian in the target image data belongs; The step of performing high-order pooling on the preset number of second features to obtain first high-order features with the same index and second high-order features with cross-indexes includes: Calculating the preset number of second features by a first high-order algorithm to obtain first high-order features with the same index; Obtaining a preset number of third features by performing channel shuffling processing on the preset number of second features; Calculating the preset number of third features by a second high-order algorithm to obtain cross-indexed second high-order features; Wherein, the first high-order algorithm is: Among them, x s is the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number; Wherein, the second high-order algorithm is: Among them, x c is the second highest-order feature of the cross-index, ⊙ represents the Hadamard product, represents the i-th third feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number; Among them, the method of obtaining a preset number of third features by performing channel shuffling processing on the preset number of second features includes: obtaining query features and key features corresponding to each second feature by performing convolution processing on the preset number of second features respectively; obtaining interactive attention features corresponding to the same second feature by performing inner product processing on the query features and the corresponding key features corresponding to the same second feature; and obtaining the third feature corresponding to the same second feature by performing matrix multiplication on the interactive attention features and the same second feature.

2. The method for pedestrian re-identification according to claim 1, It is characterized in that The method of obtaining target image data by preprocessing the image data includes: Processing the image data into a target size to obtain first image data; The target image data is obtained by performing normalization processing on the first image data.

3. The method for pedestrian re-identification according to claim 1, It is characterized in that The step of performing convolution processing on the first feature by using the pedestrian re-identification model to obtain a preset number of second features includes: Inputting the first features into a preset number of convolutional layers for processing respectively to obtain a preset number of first sub-features; A preset number of second features are obtained by performing batch normalization processing on the preset number of first sub-features respectively.

4. The method for pedestrian re-identification according to claim 1, It is characterized in that The step of calculating the preset number of third features by using a second high-order algorithm to obtain cross-indexed second high-order features includes: Calculating the preset number of third features a preset number of times by a second high-order algorithm to obtain a corresponding preset number of cross-indexed second high-order features; The preset number of cross-indexed second high-order features are fused through a preset algorithm to obtain final cross-indexed second high-order features.

5. The method for pedestrian re-identification according to claim 1, It is characterized in that The training of the pedestrian re-identification model includes: Performing random size processing on the training sample image to obtain a first training sample image; Processing the first training sample image into a target size to obtain a second training sample image; Obtaining target training sample data by normalizing the second training sample image; Construct an initial person re-identification model, set model parameters and corresponding update algorithms, and obtain the first person re-identification model; Training the first person re-identification model using the target training sample data, optimizing model parameters using a joint loss function, and determining whether the initial person re-identification model is trained properly; If the training is qualified, a second person re-identification model is obtained, and the second person re-identification model is tested by using a test sample image to obtain a test result; When the test result indicates that the recognition of the test sample image is correct, the second person re-recognition model is determined as a finally applicable person re-recognition model.

6. The method for pedestrian re-identification according to claim 5, It is characterized in that Set model parameters and corresponding update algorithms, including: Initializing the network weight parameters of the initial person re-identification model through the ImageNet pre-training model; The network weight parameters are updated during the training process using the stochastic gradient descent algorithm.

7. The method for pedestrian re-identification according to claim 5, It is characterized in that The joint loss function includes: a triplet loss function and an attention regularization function.

8. A system for pedestrian re-identification, It is characterized in that The system comprises: A preprocessing module, used for obtaining target image data by preprocessing the image data; A feature extraction module, used for inputting the target image data into a backbone network of a person re-identification model for feature extraction to obtain a first feature; A convolution processing module, used for performing convolution processing on the first feature through the pedestrian re-identification model to obtain a preset number of second features; A high-order pooling module, used to obtain first high-order features with the same index and second high-order features with cross-indexes by performing high-order pooling processing on the preset number of second features; A feature fusion module, used for obtaining a target high-order feature by fusing the first high-order feature and the second high-order feature; A similarity calculation module, used for calculating the similarity between the target high-order feature and each candidate high-order feature to obtain a calculation result; A pedestrian re-identification module, used to determine the target pedestrian to which the pedestrian in the target image data belongs according to the calculation result; Wherein, the high-order pooling module includes: A first high-order pooling module, used to calculate the preset number of second features by using a first high-order algorithm to obtain first high-order features with the same index; a channel disorder processing module, used for obtaining a preset number of third features by performing channel disorder processing on the preset number of second features; A second high-order pooling module, used to calculate the preset number of third features by a second high-order algorithm to obtain cross-indexed second high-order features; Wherein, the first high-order algorithm is: Among them, x s is the first high-order feature with the same index, ⊙ represents the Hadamard product, represents the i-th second feature X i In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number; Wherein, the second high-order algorithm is: Among them, x c is the second highest-order feature of the cross-index, ⊙ represents the Hadamard product, represents the i-th third feature In p x Component features of position, V x represents the location set of all component features, |·| represents the set capacity, and n is the preset number; Wherein, the channel disorder processing module includes: A first convolution module is used to obtain query features and key features corresponding to each second feature by performing convolution processing on the preset number of second features respectively; An inner product module, used for obtaining an interactive attention feature corresponding to the same second feature by performing inner product processing on a query feature and a corresponding key feature corresponding to the same second feature; A matrix multiplication module is used to obtain a third feature corresponding to the same second feature by performing matrix multiplication on the interactive attention feature and the same second feature.