A dynamic training method and system for person re-identification under progressive multi-loss function constraints

By adopting a dynamic training method under the constraints of the progressive multi-loss function in pedestrian re-identification technology, the poor recognition performance problem caused by large posture changes and similar clothing is solved, and a higher recognition accuracy and lower error detection and missed detection rate are achieved.

CN113609920BActive Publication Date: 2025-05-13HANGZHOU YINGGEZHIDA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202110785682.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-12
Publication Date
2025-05-13
Estimated Expiration
2041-07-12

AI Technical Summary

Technical Problem

The existing pedestrian recognition technology has poor model recognition performance when faced with large posture changes and similar to clothing, and is prone to missed detection.

Method used

The pedestrian recognition dynamic training method is adopted under the constraint of the progressive multi-loss function. Through standardized alignment, multi-scale feature extraction and progressive multi-loss joint constraint dynamic training, the intra-class distance and inter-class distance are regulated, and the model feature extraction ability is improved.

Benefits of technology

It significantly improves the recognition performance of the model, reduces false detection and missed detection rates, and optimizes the model effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113609920B_ABST
    Figure CN113609920B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of deep learning, and discloses a dynamic training method and system for pedestrian re-identification under the constraints of progressive multi-loss functions; the position of the pedestrian cropping frame is determined by key point detection technology, and the pedestrian parts are aligned by padding; the network structure is designed to extract detailed features of different sizes; the extracted feature vectors are dynamically trained by a progressive multi-loss joint constraint method. The similarity between each image block is compared to achieve more accurate and more effective feature comparison; a multi-scale feature extraction module is designed to capture detailed features of different sizes and increase the network feature extraction capability; a progressive multi-loss joint constraint method is used for dynamic training, which has strong controllability and high flexibility; according to different training stages, the intra-class distance and the inter-class distance can be gradually controlled to continuously tend to the ideal state, and finally the model feature extraction capability is improved and the model effect is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a dynamic training method and system for pedestrian re-identification based on the constraints of a progressive joint loss function. Background Art

[0002] Pedestrian re-identification, also known as pedestrian re-ID, is a technology that uses computer vision technology to determine whether there is a specific pedestrian in an image or video sequence. This technology is very important in the field of vision in surveillance equipment.

[0003] Although current facial recognition technology is very mature and can identify a person through facial comparison search and identification, it also has some limitations. For example, facial recognition will not be able to identify pedestrians in special circumstances such as wearing masks or facing away from the surveillance camera.

[0004] Pedestrian re-identification technology can retrieve images of the pedestrian under different devices given a monitored pedestrian image. Pedestrian re-identification technology can continuously track pedestrians whose faces cannot be clearly photographed across cameras, enhance the spatiotemporal continuity of data, and collaborate with face recognition technology in security and missing children scenarios to assist in efficient incident handling. However, current research on pedestrian re-identification also faces challenges such as large posture changes and overly similar clothing. How to improve model recognition performance and reduce false detections and missed detections has become an important research branch in the industry.

[0005] The patent name is: A pedestrian re-identification method and system, the application number is: CN201910672444.2, the application date is: 2019-07-24, the patent application discloses that the pedestrian analysis network is used to analyze the input image, and the fine-grained features of the pedestrians in the input image are extracted; the fine-grained features are fused with the pedestrian features of the input image output by the convolution layer of the pedestrian re-identification network model; according to the fused pedestrian features, the pedestrians in the input image are identified. The patent uses the commonly used Resnet50 in the industry as the skeleton network, while the present invention uses a self-designed parallel multi-branch multi-scale feature extraction structure as the backbone network, which has stronger detail feature mining capabilities. The loss function and model training method involved in the present invention are completely different from those of the patent. The patent uses ordinary triple loss function and cross entropy loss function, and uses a general training scheme. However, the present invention takes into account the optimization difficulty and inconsistency of directions for each category, sets different control factors for each category, and adopts a progressive training strategy during the training process.

[0006] The patent name is: A pedestrian re-identification method that enhances local feature learning by combining multiple loss dynamic training strategies, the application number is: CN2020109348839, the application date is: 2020-09-08. The patent application discloses a method of using the idea of ​​component alignment for feature matching, and using a self-attention mechanism network to extract non-pedestrian local features; the patent emphasizes that the local features are mainly items carried by pedestrians, and the global features of pedestrians and the local features are fused for pedestrian re-identification. In terms of training methods, dynamic training cross entropy and triplet loss function methods are used to optimize model parameters. The difference is that the alignment method of the present invention is different from that of the patent, and the network structure used is completely different. The dynamic training method involved in the patent is to add an average moving factor to the average loss value of the current two loss functions, and control the constraint strength by controlling the average moving factor. The dynamic training process in the present invention is more precise and detailed. Summary of the invention

[0007] In view of the problems in the prior art of pedestrian re-identification, such as large posture changes and too similar clothing, poor model recognition performance, and prone to false detection and missed detection, the present invention provides a dynamic training method and system for pedestrian re-identification under progressive multi-loss function constraints.

[0008] In order to solve the above technical problems, the present invention is solved by the following technical solutions:

[0009] A dynamic training method for pedestrian re-identification under progressive multi-loss function constraints, the method includes

[0010] Standardized alignment: determine the position of the pedestrian cropping frame through key point detection technology, and align pedestrian parts through padding;

[0011] Multi-scale feature extraction: design network structure to extract detailed features of different sizes;

[0012] Progressive multi-loss joint constraint dynamic training: The extracted feature vectors are dynamically trained through the progressive multi-loss joint constraint method.

[0013] Through standardized alignment, the similarity between each image block is compared to achieve more accurate and effective feature comparison; the image recognition problem is upgraded to an instance recognition problem, and a multi-scale feature extraction module is designed to capture detailed features of different sizes and increase the network feature extraction capability; the progressive multi-loss joint constraint method is used for dynamic training, which has strong controllability and high flexibility; according to different training stages, the intra-class distance and the inter-class distance can be gradually controlled to continuously tend to the ideal state, ultimately improving the model feature extraction capability and optimizing the model effect.

[0014] As a preferred method, the progressive multi-loss joint constraint dynamic training method includes the FixTriplet and HAAM joint loss functions. The constraint interval is enlarged by the FixTriplet joint loss function, requiring that the maximum intra-class distance should always be smaller than the inter-class distance of any two pedestrians. The loss function has a larger fault tolerance space. If an input pair with extremely small inter-class distance or extremely large intra-class distance appears, the input is considered to be an outlier and the calculation result is optimized.

[0015] As a preferred approach, progressive multi-loss joint constrained dynamic training,

[0016] The first step is to calculate the maximum intra-class distance in each identity A, the identity C corresponding to identity B with the minimum inter-class distance, and randomly sample from A, B, and C. By pre-calculating the identity B with the minimum inter-class distance corresponding to each identity A, and randomly sampling from A and B, a valid triple input is constructed.

[0017] , that is, FixTriplet loss function L FTP ;

[0018]

[0019] In the formula, A1 and A2 represent two different images of the same pedestrian, B and C represent two pedestrian images with different identities, and the condition ABC contains at least two pedestrians; [·] + represents the function max(·,0), m represents the hyperparameter of the control distance, MA represents the maximum intra-class distance, MI represents the minimum inter-class distance, ε and η represent the correction values, and s1 and s2 represent the critical outliers, respectively;

[0020] The second step is to use the loss function HAAM loss function L HAAM The features are mapped onto a hypersphere.

[0021]

[0022] In the formula, s represents the radius of the HAAM loss hypersphere, N represents the number of input images in each iteration, and y i Indicates the index value corresponding to the input label in each iteration, and cosθ j (j≠y i ) represent the intra-class similarity and inter-class similarity respectively, m i represents the distance control hyperparameter, k represents the total number of classifications in the training set, and a and b are control parameters;

[0023] The third step is to use the loss function L HAAM and L FTP Get the final loss value lall .

[0024] Preferably, the key point detection technology determines that the feature vector includes the human head, shoulder joint, and hip joint.

[0025] As a preference, standardized alignment,

[0026] Determine the position of pedestrians and lock the position of pedestrians using key point detection technology;

[0027] Conversion of cropping frame coordinate information, converting the key point coordinates of the pedestrian position into cropping frame coordinate information,

[0028] Acquisition of image data, by acquiring image data within the cropping frame coordinate information;

[0029] Image data aspect ratio judgment: For the acquired image data, it is judged whether the image aspect ratio is a standard value. If it is not a standard value, it is aligned by padding.

[0030] Preferably, the aspect ratios include head-to-shoulder ratio, shoulder-to-hip ratio, and lower limb size ratio.

[0031] As a preference, multi-scale feature extraction

[0032] The feature vector is obtained by 1*1 convolution dimensionality reduction, and the obtained feature vector is input into the pyramid-like structure. The features obtained by three convolution groups with different depths are used through the spatial attention mechanism to obtain the feature vectors of different parts;

[0033] After four times of multi-scale feature extraction, the network output F is input into two branches at the same time. The first branch divides F into three parts unevenly, and the second branch divides F into two parts equally to assist in feature alignment.

[0034] The five branches are subjected to global average pooling and convolution operations to obtain feature matrices, each of which contains N feature vectors of length 512. By calculating the five feature vectors in the above two branches and comparing the similarity between each feature block, more accurate and effective feature comparison can be achieved.

[0035] A dynamic training system for person re-identification under progressive multi-loss function constraints, including a standardized alignment module, a multi-scale feature extraction module and a dynamic training module;

[0036] The standardized alignment module is used to align feature vectors. The feature vectors are determined by key point detection technology and aligned by padding.

[0037] Multi-scale feature extraction module, used to extract detail features, combined with the network structure to extract detail features of different sizes;

[0038] Dynamic training module; used to obtain loss values ​​and dynamically train the extracted feature vectors through a progressive multi-loss joint constraint method.

[0039] Preferably, the progressive multi-loss joint constraint method includes FixTriplet and HAAM joint loss functions.

[0040] The present invention has significant technical effects due to the adoption of the above technical solutions: the present invention adopts a self-designed network structure, uses convolution structures of different depths to extract effective features of different scales, and performs strong data standardization operations to achieve a great degree of part alignment; a new joint constraint dynamic training strategy is designed, which can achieve progressive adjustment of training difficulty, optimize training results, and improve model accuracy.

[0041] Through standardized alignment, the similarity between each image block is compared to achieve more accurate and effective feature comparison; the image recognition problem is upgraded to an instance recognition problem, and a multi-scale feature extraction module is designed to capture detailed features of different scales and increase the network feature extraction capability; the progressive multi-loss joint constraint method is used for dynamic training, which has strong controllability and high flexibility; according to different training stages, the intra-class distance and the inter-class distance can be gradually controlled to continuously tend to the ideal state, ultimately improving the model feature extraction capability and optimizing the model effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flow chart of the present invention.

[0043] Figure 2 2 is a diagram of the pedestrian re-identification network structure of the present invention.

[0044] Figure 3 This is a network structure diagram of the multi-scale feature extraction module in the present invention. DETAILED DESCRIPTION

[0045] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0046] Example 1

[0047] A dynamic training method for pedestrian re-identification under progressive multi-loss function constraints, the method includes

[0048] Standardized alignment: determine the position of the pedestrian cropping frame through key point detection technology, and align pedestrian parts through padding;

[0049] Multi-scale feature extraction: design network structure to extract detailed features of different sizes;

[0050] Progressive multi-loss joint constraint dynamic training: The extracted feature vectors are dynamically trained through the progressive multi-loss joint constraint method.

[0051] Through standardized alignment, the similarity between each image block is compared to achieve more accurate and effective feature comparison; the image recognition problem is upgraded to an instance recognition problem, and a multi-scale feature extraction module is designed to capture detailed features of different sizes and increase the network feature extraction capability; the progressive multi-loss joint constraint method is used for dynamic training, which has strong controllability and high flexibility; according to different training stages, the intra-class distance and the inter-class distance can be gradually controlled to continuously tend to the ideal state, ultimately improving the model feature extraction capability and optimizing the model effect.

[0052] The progressive multi-loss joint constraint dynamic training method mainly relies on the FixTriplet and HAAM joint loss functions. This joint loss has a certain fault tolerance space, can handle outliers of erroneous data, has a dynamic adjustment mechanism for training difficulty, is highly flexible, and has strong adaptability, and has stronger guidance for the model training direction.

[0053] Progressive multi-loss joint constrained dynamic training,

[0054] The first step is to calculate the maximum intra-class distance of each identity A, the identity B corresponding to the minimum inter-class distance identity C, and randomly sample from A, B and C to construct a valid triplet input, that is, the FixTriplet loss function L FTP ;

[0055]

[0056] In formula 1, A1 and A2 represent two different images of the same pedestrian, B and C represent two images of pedestrians with different identities, and the condition ABC contains at least two pedestrians; [·] + represents the function max(·,0), m represents the hyperparameter of the control distance, MA represents the maximum intra-class distance, MI represents the minimum inter-class distance, ε and η represent the correction values, and s1 and s2 represent the critical outliers, respectively;

[0057] The second step is to map the features onto a hypersphere through the loss function HAAM.

[0058]

[0059] In Formula 2, s represents the radius of the HAAM loss hypersphere, N represents the number of input images in each iteration, and y i Indicates the index value corresponding to the input label in each iteration, and cosθ j (j≠y i) represent the intra-class similarity and inter-class similarity respectively, m i represents the distance control hyperparameter, k represents the total number of classifications in the training set, and a and b are control parameters;

[0060] The third step is to use the loss function L HAAM and L FTP Get the final loss value l all .

[0061] Key point detection technology determines that the feature vector includes the human head, shoulder joint, and hip joint.

[0062] Standardized alignment, pedestrian position determination, and key point detection technology are used to lock the pedestrian position;

[0063] Conversion of cropping frame coordinate information, converting the key point coordinates of the pedestrian position into cropping frame coordinate information,

[0064] Acquisition of image data, by acquiring image data within the cropping frame coordinate information;

[0065] Image data aspect ratio judgment: For the acquired image data, it is judged whether the image aspect ratio is a standard value. If it is not a standard value, it is aligned by padding.

[0066] Aspect ratios include the head-to-shoulder ratio, the shoulder-to-hip ratio, and the lower limb size ratio.

[0067] Multi-scale feature extraction

[0068] The features obtained by 1*1 convolution dimensionality reduction are input into the pyramid-like structure, and then pass through three convolution groups of different depths. The spatial attention mechanism fuses the effective information at different positions to obtain high-quality features;

[0069] After four times of multi-scale feature extraction, the network output F is input into two branches at the same time. The first branch divides F into three parts unevenly, and the second branch divides F into two parts equally to assist in feature alignment.

[0070] The five branches are subjected to global average pooling and convolution operations to obtain feature matrices, each of which contains N feature vectors of length 512.

[0071] Example 2

[0072] A dynamic training system for person re-identification under progressive multi-loss function constraints, a standardized alignment module for aligning feature vectors, determining feature vectors through key point detection technology, and aligning feature vectors through padding;

[0073] Multi-scale feature extraction is used to extract size detail features and extract detail features of different sizes in combination with the network structure;

[0074] Dynamic training module: used to obtain loss values ​​and dynamically train the extracted feature vectors through a progressive multi-loss joint constraint model. The progressive multi-loss joint constraint method includes FixTriplet and HAAM joint loss functions.

[0075] Example 3

[0076] On the basis of the above-mentioned embodiment, the video data is monitored by the human body key point positioning technology, and the cropping frame is obtained according to the cropping frame coordinates and key point coordinates conversion formula to obtain preliminary data.

[0077] Then the aspect ratio of a single image is calculated, and padding is added to images whose aspect ratio is not equal to the standard value. The above process realizes the standardization of data and lays a good foundation for subsequent feature comparison.

[0078] In each iteration, N images are input and sent to the multi-scale feature extraction module. There are four consecutive multi-scale feature extraction modules in the entire network. The following will take a single multi-scale feature extraction module as an example to explain the technical principle in detail. Figure 3 As shown in the figure, the structure refers to the ResNeXt network topology, designs a parallel convolution structure, increases the network width to improve the model's feature extraction capabilities, and draws on the feature pyramid structure to build a rich receptive field by stacking convolutional layers of different depths. This effectively extracts target feature information of different sizes in the image, and this part is called a pyramid-like structure.

[0079] First, 1*1 convolution is used for dimensionality reduction, and then the features enter the pyramid-like structure. The features obtained through three different depth convolution groups are used through the spatial attention mechanism to obtain the significant features of different parts. After four multi-scale feature extractions, the network output F is input into two branches at the same time. The first branch divides F into three parts unevenly, and the second branch divides F into two parts equally to assist in feature alignment. Subsequently, the five branches are subjected to global average pooling and convolution operations to obtain feature matrices, each of which contains N feature vectors of length 512.

[0080] The whole process of extracting features from the network is abstracted as f(x), where x represents the input data. These features are then input into the FixTriplet loss function and the HAAM loss function to calculate the results.

[0081] Calculate FixTriplet loss L FTP :

[0082]

[0083]

[0084] According to the feature extraction results of the neural network: after inputting the pedestrian images A1 and A2 of identity A and the pedestrian images of identity B and C, the corresponding f(A1), f(A2), f(B), f(C) are obtained, and d(f(A1), f(A2)), d(f(B), f(C)) are calculated according to the Euclidean distance calculation formula. Design correction values ​​ε and k, and then determine whether MA is greater than s1 and whether MI is less than s2. If so, perform corresponding processing according to the formula and update to the corresponding correction value. Calculate L FTP , and finally calculate the average value to obtain l ftp .

[0085]

[0086] Calculate L corresponding to the five branches HAAM . When calculating the HAAM loss function, multi-stage constraint strength control is implemented. In the formula, T represents the number of iterations, ep1 and ep2 represent the set training stage points, and ep3 represents the total number of iterations. p and q represent the rate of change, u and v represent the offset, which constitute two sets of linear change functions respectively. In the initial stage of training, the model's feature extraction ability is poor, and the model parameters change rapidly. When comparing feature similarity, the calculation results are relatively of no reference value. Therefore, in the early stage of training, when the number of iterations T is less than the set initial number of iterations ep1, we set the values ​​of a and b to 0. After several cycles of iteration, the model has a certain feature extraction ability. At this time, the values ​​a and b are increased, and a and b are passed into two linear functions to achieve the purpose of progressive constrained dynamic training. Finally, to ensure stable convergence of the model, the parameters of a and b are fixed.

[0087]

[0088]

[0089] The number of all pedestrian identity categories is k. The model calculates the similarity between the extracted feature X and the k pedestrian class center vectors W, finds the most similar one, and then determines the pedestrian identity category. The class center vector W is initialized to a size of k×512. The training purpose is to guide the distance between these class centers to be as large as possible through the loss function constraint. i Indicates the index value corresponding to the position of the pedestrian identity label in each iteration, j (j≠y i ) represents the index value of the non-corresponding position. In the HAAM loss function calculation process, first, calculate W, X normalized result W norm , X norm The matrix product of k =Wnorm *X norm Then the obtained Cos k Can be divided into two subsets and Cos j , Represents the similarity between pedestrian features and other features in the same category. Similarly, Cos j represents the similarity between pedestrian images of different categories. Next, calculate Calculate the different m corresponding to each identity category in each iteration i , calculate L HAAM .

[0090]

[0091]

[0092] Calculate the average HAAM loss result l HAAM .

[0093]

[0094] Combining the above two loss results, we get the final loss value l all =2*l HAAM +l FTP As the model gradually converges, the model feature extraction capability improves, and the parameter m i The constraints on the model are continuously increased, guiding the distance between classes to continuously approach the ideal value, thus completing the dynamic training of the model.

Claims

1. A dynamic training method for person re-identification under progressive multi-loss function constraints, characterized in that: Methods include: Standardized alignment: determine the position of the pedestrian cropping frame through key point detection technology, and align pedestrian parts through padding; Multi-scale feature extraction: design network structure to extract detailed features of different sizes; Progressive multi-loss joint constrained dynamic training; The extracted feature vectors are dynamically trained through a progressive multi-loss joint constraint method; Progressive multi-loss joint constrained dynamic training, including: The first step is to calculate the maximum intra-class distance of each identity A, the identity B corresponding to the minimum inter-class distance identity C, and randomly sample from A, B and C to construct a valid triplet input, that is, the FixTriplet loss function L FTP ; In the formula, A1 and A2 represent two different images of the same pedestrian, B and C represent two pedestrian images with different identities, and the condition ABC contains at least two pedestrians; [·] + represents the function max(·,0), m represents the hyperparameter of the control distance, ε and η represent the correction values, s1 and s2 represent the critical outliers, respectively; The second step is to use the HAAM loss function L HAAM Mapping the features onto a hypersphere, In the formula, s represents the radius of the HAAM loss hypersphere, N represents the number of input images in each iteration, and y i Indicates the index value corresponding to the input label in each iteration, and cosθ j Respectively represent the intra-class similarity and inter-class similarity, m i represents the distance control hyperparameter, k represents the total number of categories in the training set, a and b are control parameters, j≠y i ; The third step is to use the loss function L HAAM and L FTP Get the final loss value l all .

2. The method for dynamic training of person re-identification under progressive multi-loss function constraints according to claim 1, characterized in that: Key point detection technology determines that the feature vector includes the human head, shoulder joint, and hip joint.

3. The dynamic training method for person re-identification under progressive multi-loss function constraints according to claim 1, characterized in that: Standardized alignment, including: Determine the position of pedestrians and lock the position of pedestrians using key point detection technology; Conversion of cropping frame coordinate information, converting the key point coordinates of the pedestrian position into cropping frame coordinate information, Acquisition of image data, by acquiring image data within the cropping frame coordinate information; Image data aspect ratio judgment: For the acquired image data, it is judged whether the image aspect ratio is a standard value. If it is not a standard value, it is aligned by padding.

4. The dynamic training method for person re-identification under progressive multi-loss function constraints according to claim 1, characterized in that: Aspect ratios include the head-to-shoulder ratio, the shoulder-to-hip ratio, and the lower limb size ratio.

5. The dynamic training method for person re-identification under progressive multi-loss function constraints according to claim 1, characterized in that: Multi-scale feature extraction, including: The feature vector is obtained by 1*1 convolution dimensionality reduction, and the obtained feature vector is input into the pyramid-like structure. The features obtained by three convolution groups with different depths are used through the spatial attention mechanism to obtain the feature vectors of different parts; After four times of multi-scale feature extraction, the network output F is input into two branches at the same time. The first branch divides F into three parts unevenly, and the second branch divides F into two parts equally to assist in feature alignment. The five branches are subjected to global average pooling and convolution operations to obtain feature matrices, each of which contains N feature vectors of length 512.

6. A dynamic training system for person re-identification under progressive multi-loss function constraints, used to execute the method according to any one of claims 1 to 5, characterized in that: include: The standardized alignment module is used to align feature vectors. The feature vectors are determined by key point detection technology and aligned by padding. Multi-scale feature extraction is used to extract size detail features and extract detail features of different sizes in combination with the network structure; Dynamic training module; used to obtain loss values ​​and dynamically train the extracted feature vectors through a progressive multi-loss joint constraint model.

Citation Information

Patent Citations

  • Pedestrian re-identification method and system

    CN110378301A

  • Pedestrian re-identification method using attitude information to design multi-loss function

    CN107832672A

  • Pedestrian re-identification method based on global feature and local feature splicing

    CN111666843A