Cross-view pedestrian re-identification method for identity information matrix
By using low-rank decomposition learning dictionary pairs and identity information matrix in pedestrian re-identification technology, the problem of difficult pedestrian feature information is solved in different perspectives, and higher recognition accuracy and stability are achieved.
Patent Information
- Application Number
- CN202311564040.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-23
AI Technical Summary
Existing pedestrian recognition technology is difficult to obtain discriminant feature information from different camera perspectives, and pedestrian feature information is prone to change during training, affecting the accuracy of subsequent matching.
Through low-rank decomposition learning dictionary pairs, the domain general dictionary and domain invariant dictionary are separated, and the identity information matrix is combined, and the classifier is trained to improve recognition performance.
It effectively overcomes the domain offset problem at different perspectives, improves the accuracy and stability of pedestrian re-identification, and improves the recognition performance.
Smart Images

Figure CN120032389A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a cross-view pedestrian re-identification method oriented to an identity information matrix, and belongs to the field of pattern recognition. Background Art
[0002] Pedestrian re-identification is a technology that uses computers to process the obtained pictures and videos with visual technology, and then compares them with other pictures and videos to find out whether there are matching pedestrians. In the era of information explosion, pedestrian re-identification has broad application prospects, but it also faces huge challenges in real society. After continuous exploration by researchers, the research directions of pedestrian re-identification algorithms at home and abroad can be roughly divided into four: feature extraction, distance measurement, dictionary learning and deep learning. In real applications, images are usually obtained under surveillance cameras with a large field of view, which are not controlled by the environment, resulting in relatively low resolution, which makes it difficult for people to obtain discriminative feature clues. Therefore, obtaining discriminative information under different camera viewing angles has become the primary problem of pedestrian re-identification. At the same time, the pedestrian feature information obtained by the camera will also be changed after multiple debugging during the training process, and the gap with the original pedestrian information is too large, which has different degrees of negative impact on the subsequent pedestrian re-identification matching. The present invention discloses a pedestrian re-identification algorithm for identity information matrix to solve the two problems described above at the same time. Specifically, the identity information is embedded into the de-noised pedestrian appearance features to improve the recognition performance of the algorithm. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a cross-view pedestrian re-identification method oriented to an identity information matrix, which combines pedestrian identity information on the basis of removing image domain information to overcome the domain offset problem caused by changes in background, lighting, etc.
[0004] The technical solution of the present invention is: a cross-view pedestrian re-identification method for identity information matrix, comprising the following steps:
[0005] 1) A dictionary pair is learned through low-rank decomposition. One dictionary is used to represent the interference information other than pedestrians in the image, that is, the domain-general dictionary, and the other dictionary represents the pedestrian feature information after removing the interference information, that is, the domain-invariant dictionary;
[0006] 2) training the discriminant promotion items of the dictionary to make the two different dictionaries independent of each other;
[0007] 3) Training the discriminative promotion term of pedestrian encoding coefficients, forcing the encoding coefficients of pedestrian images with the same viewpoint to have strong similarity;
[0008] 4) The pedestrian encoding coefficients obtained by training the pedestrian feature information after removing the domain information are combined with the identity information to learn the classifier. At the same time, the consistent matrix and irrelevant matrix of the identity information are introduced to give the classifier a higher discrimination ability;
[0009] 5) Determine the overall objective function of cross-view person re-identification for the identity information matrix;
[0010] 6) Solving the variables to be updated in the overall objective function;
[0011] 7) Design a pedestrian matching solution based on the model that introduces the identity information matrix after removing the domain information.
[0012] Specifically, the dictionary pair of step 1) includes:
[0013] use represents the training set obtained under camera a, represents the training set obtained under camera b, where m represents the feature dimension and n represents the number of pedestrians. D is the domain information dictionary shared by all cameras, that is, the domain universal dictionary, D t is the pedestrian feature dictionary under a single camera, that is, the domain-invariant dictionary. The domain-general dictionary and the domain-invariant dictionary are integrated together. At this time, the dictionary learning algorithm is shown in formula (1):
[0014]
[0015] In the formula, S a and S b Corresponding training set X on dictionary D a and X b The encoding coefficient matrix, S ta and S tb Corresponding dictionary D t Pedestrian coding coefficients under different viewing angles. The connection between the pedestrian training set of camera a and the domain general dictionary D is established. Same here. and It means to separate the information of domain-general dictionary from the information of domain-invariant dictionary. t ,S a ,S b ,S ta ,S tb ) is the discriminant promotion term of the dictionary Ψ(D,D t ) and the discriminant promotion term Γ(S a ,S b ,S ta ,S tb ). Represents the F-norm.
[0016] Specifically, the dictionary discrimination promotion items of step 2) include:
[0017] The proposed dictionary discrimination boosting term is:
[0018]
[0019] In the formula, ||D|| * To solve the nuclear norm of the dictionary D, that is, the sum of the singular values of the dictionary D. At the same time, the domain-universal dictionary and the domain-invariant dictionary have different spatial forms. To make the two dictionaries independent of each other, T represents transposition. 1 and α 2 It is a scalar parameter representing the weight information of the contained items.
[0020] Specifically, the discrimination promotion item of the pedestrian coding coefficient in step 3) includes:
[0021] The proposed discriminant promotion term of the pedestrian coding coefficient is:
[0022] Γ(S a ,S b ,S ta ,S tb )=α 3 {||S a || 2,1 +||S b || 2,1}+α 4 {||S ta || 1 +||S tb || 1} (3)
[0023] In the formula, ||S a || 2,1 and ||S b || 2,1 for l 2,1 norm, used to regularize the matrix, specifically, ||S ta || 1 and ||S tb || 1 for l 1 norm, used to improve the coding coefficient S ta and S tb The sparsity of α 3 and α 4 It is a scalar parameter representing the weight information of the contained items.
[0024] Specifically, the classifier of step 4) includes:
[0025] Combine the pedestrian encoding coefficients obtained through pedestrian feature information training with identity information to train the classifier W a and W b , realizing the association between the pedestrian appearance features and identity information under two different perspectives a and b. At the same time, the identity consistency matrix and the irrelevant matrix are introduced on the basis of the original model to give the classifier a higher discrimination ability.
[0026] The proposed classifier is:
[0027]
[0028] In the formula, They correspond to the pedestrian identity information matrices under cameras a and b respectively, and k represents the pedestrian category. Represents the identity information matrix between different perspectives. When the visual feature j obtained by camera a and the visual feature k obtained by camera b come from the same person, E(j,k) = 1, otherwise E(j,k) = 0. 5 and α 6 Scalar parameter, representing the weight information of the contained items. The understanding of is: if the training picture is a picture of the same pedestrian under the perspective a, the value obtained is 1; if the training picture is a picture of different pedestrians, the value obtained is 0. Similarly, It represents the b perspective, which has the same function as the a perspective and will not be described here. It represents different viewing angles. If the pedestrian is the same in the two viewing angles, the value is 1. If the pedestrian is different in the two viewing angles, the value is 0.
[0029] Specifically, the overall objective function of step 5) includes:
[0030] The overall model framework of cross-view person re-identification for identity information matrix is defined as:
[0031]
[0032] In the formula, α 7 and α 8 It is a scalar parameter representing the weight information of the contained items. and Used to control model complexity and avoid overfitting. ||(W a S a ) T (W b S b )|| 1 Yes 1 The constraint of the norm can make (W a S a )T (W b S b ) into a sparse matrix, which further optimizes the computer's matrix calculations and improves the ability to distinguish identity information.
[0033] Specifically, the variable solving in step 6) includes:
[0034] The variables D, D in formula (5) t ,S a ,S b ,W a ,W b ,S ta ,S tb It cannot be solved directly. Use the iterative shrinkage method to solve different variables separately. The solution of each variable is as follows:
[0035] Fixed D,D t ,S a ,S b ,W b ,S ta ,S tb Update variable W a , then the equation (5) about W a In the objective function, W a The terms can be expressed as:
[0036]
[0037] Formula (6) requires the introduction of an intermediate variable A. The algorithm is as follows:
[0038]
[0039] In formula (7), A is a typical l 1 The norm minimization problem can be solved by using the iterative shrinkage algorithm. After determining A, it is still necessary to introduce an intermediate variable M = W a S ta , then you can a Update. The algorithm is as follows:
[0040]
[0041] The variable M in formula (8) can be directly derived to obtain its solution:
[0042]
[0043] In formula (9), The identity matrix. After the variable M is calculated, W a It can also be directly derived to obtain its solution:
[0044]
[0045] in, is the identity matrix, and d is the size of the dictionary D.
[0046] Using the same method, fix D,D t ,S a ,S b ,W a ,S ta ,S tb Update variable W b , then the equation (5) about W b The objective function can be expressed as:
[0047]
[0048] According to the method in formula (7), an intermediate variable B is introduced to solve the typical l 1 Norm minimization problem. And according to the method of formula (8), an intermediate variable N = W is introduced b S tb After determining B and N, we can calculate W b Direct derivation to obtain its analytical source is:
[0049]
[0050]
[0051] in, is the identity matrix, is the identity matrix, d t It is a dictionary D t size.
[0052] Fixed D,D t ,S b ,W a ,W b ,S ta ,S tb No change, update S a Variable, formula (5) about S a The objective function can be expressed as:
[0053]
[0054] In formula (14), ||S a || 2,1 for l 2,1 Paradigm, find S a The parsed source is:
[0055]
[0056] Among them, Λ 1 The diagonal sparse matrix, S a j Indicates S a The jth column of .
[0057] Use the same method to update the variable S b , fixed D,D t ,S a ,W a ,W b ,S ta ,S tb unchanged, the formula (5) about S b The objective function can be expressed as:
[0058]
[0059] Solve l according to the method in formula (14) 2,1 Paradigm problem, that is, S b Direct derivation yields the following solution:
[0060]
[0061] Among them, Λ 2 for The diagonal sparse matrix, S b j Indicates S b The jth column of .
[0062] Fixed D,D t ,S a ,S b ,W a ,W b ,S tb No change, update S ta Variable, formula (5) about S ta The objective function can be expressed as:
[0063]
[0064] Two intermediate variables C and G need to be introduced in formula (18) to solve l 1 Paradigm problem, equation (18) can be changed to:
[0065]
[0066] S in formula (19) ta It is still not directly solvable, and an intermediate variable O=W needs to be introduced a S ta , rewritten as:
[0067]
[0068] The analytical source of O in formula (20) can be directly obtained as:
[0069]
[0070] in, is the identity matrix.
[0071] After obtaining O, you can ask for S ta , containing S ta The items are:
[0072]
[0073] Therefore, S can be obtained by equation (22): ta Direct derivation yields:
[0074]
[0075] in, is the identity matrix.
[0076] Use the same method to update the variable S tb , fixed D,D t ,S a ,S b ,W a ,W b ,S ta unchanged, the formula (5) about S tb The objective function can be expressed as:
[0077]
[0078] According to the solution method of formula (18), the l in formula (24) can be solved 1 Paradigm problem, formula (24) can be written as:
[0079]
[0080] Formula (25) can be used to calculate S tb Directly solve to obtain its analytical source:
[0081]
[0082] in, is the identity matrix.
[0083] Fixed D t ,S a ,S b ,W a ,W b,S ta ,S tb Keep the variable D unchanged, and update the variable D. The objective function of formula (5) with respect to D can be expressed as:
[0084]
[0085] Among them, ||D|| * is the nuclear norm, and the singular value threshold algorithm can be used to solve the nuclear norm problem. Introducing an intermediate variable Q = D, equation (27) can be rewritten as:
[0086]
[0087] The analytical source of Q can be found as:
[0088] Q=(α 2 D t D t T +I 7 )\D (29)
[0089] in, is the identity matrix.
[0090] The dictionary D can be calculated by the Lagrange dual method, and its analytical source is:
[0091]
[0092] in, is the identity matrix, Λ 3 is the diagonal matrix obtained from the Lagrange dual variables.
[0093] Finally, according to formula (27), D t Optimize and fix D,S a ,S b ,W a ,W b ,S ta ,S tb constant:
[0094]
[0095] Similarly, D in formula (31) t It can also be obtained by the Lagrange dual method, which is consistent with the solution method of D and will not be repeated here.
[0096] Specifically, the pedestrian matching scheme of step 7) includes:
[0097] In the test, we use the D,D obtained during training t ,W a ,W b , the following method is used to perform Sa ,S b ,S ta ,S tb Solve and match:
[0098]
[0099] In formula (32), X a is the test sample set, S a and S b is the coding coefficient matrix under view a and b, S ta and S tb is the encoding coefficient matrix of a specific pedestrian under view a and b. First, fix D,D t ,S ta ,S tb , for S a and S b Update, about S a and S b The terms of the objective function are:
[0100]
[0101]
[0102] S a and S b The analytical source of can be obtained by direct derivation:
[0103]
[0104]
[0105] In formulas (35) and (36), Λ 3 and Λ 4 They are and The diagonal sparse matrix formed by ||S a j || indicates S a The jth column of b j || indicates S b The jth column of S a and S b Then, the same method can be used to find S ta and S tb , whose parsing source is:
[0106]
[0107]
[0108] In formulas (37) and (38), is the identity matrix, A and B are the matrix to solve l 1 The intermediate variable of the paradigm.
[0109] The S obtained by calculating the above test sample set ta and S tb , and the classifier W obtained in the training sample set calculation a and W b To perform distance calculation:
[0110] B a =S ta W a , B b =S tb W b (39)
[0111] Using B in formula (39) a and B b Perform a matching test, assuming B a Set as the array to be tested, B b Set as the target array and calculate B by Euclidean distance a (:,i) and B b (:,j), and the obtained difference values are arranged in ascending order. The obtained order is the Rank matching rate. a (:,i) and B b The difference between (:,j) indicates B a The i-th column of B b The Euclidean distance is calculated for each column of , and the average of these gaps is calculated. The average obtained is the final matching rate.
[0112] The beneficial effects of the present invention are:
[0113] 1. In real applications, images are usually obtained from surveillance cameras with a large field of view and are not controlled by the environment, resulting in relatively low resolution, which makes it difficult to obtain discriminative feature clues. Therefore, obtaining discriminative images after removing domain information under different camera perspectives has become one of the primary issues in pedestrian re-identification.
[0114] 2. The pedestrian information in the images obtained by the camera will be changed after multiple debugging during the training process, and the difference from the original pedestrian information is too large, which has different degrees of negative impact on the subsequent pedestrian re-identification matching. Therefore, combining pedestrian features and identity information for pedestrian re-identification under different camera perspectives has become the second primary problem.
[0115] 3. The pedestrian re-identification method proposed in this invention has significantly improved recognition performance compared with other methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0116] Figure 1 is a flow chart of the present invention;
[0117] Figure 2 is a pair of pedestrian images from two camera perspectives on the PRID2011 dataset provided by an embodiment of the present invention;
[0118] Figure 3 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 1 CMC curve;
[0119] Figure 4 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 2 CMC curve;
[0120] Figure 5 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 3 CMC curve;
[0121] Figure 6 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 4 CMC curve;
[0122] Figure 7 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 5 CMC curve;
[0123] Figure 8 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 6 CMC curve;
[0124] Fig. 9 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 7 CMC curve;
[0125] Fig.10 The embodiment of the present invention provides a method for calculating the parameter α in the algorithm based on the PRID2011 dataset. 8 CMC curve;
[0126] Fig.11 It is a CMC curve for parameter d in the algorithm based on the PRID2011 data set provided by an embodiment of the present invention;
[0127] Fig.12The embodiment of the present invention provides a method for calculating the parameter d in the algorithm based on the PRID2011 dataset. t CMC curve. DETAILED DESCRIPTION
[0128] The present invention will be further described below in conjunction with the accompanying drawings and specific implementation methods.
[0129] Embodiment: In real applications, images are usually obtained from surveillance cameras with a large field of view, which are not controlled by the environment, resulting in relatively low resolution, which makes it difficult for people to obtain discriminative feature clues. Therefore, obtaining discriminative information under different camera perspectives has become the primary issue in pedestrian re-identification. At the same time, the pedestrian feature information obtained by the camera will also be changed after multiple debugging during the training process, and the difference from the original feature information is too large, which has different degrees of negative impact on subsequent pedestrian re-identification matching. Based on this idea, the present invention proposes a novel identity information matrix-oriented method for cross-perspective pedestrian re-identification.
[0130] like Figure 1 As shown, a cross-view pedestrian re-identification method for identity information matrix includes the following steps:
[0131] 1) A dictionary pair is learned through low-rank decomposition. One dictionary is used to represent the interference information other than pedestrians in the image, that is, the domain-general dictionary, and the other dictionary represents the pedestrian feature information after removing the interference information, that is, the domain-invariant dictionary;
[0132] 2) training the discriminant promotion items of the dictionary to make the two different dictionaries independent of each other;
[0133] 3) Training the discriminative promotion term of pedestrian encoding coefficients, forcing the encoding coefficients of pedestrian images with the same viewpoint to have strong similarity;
[0134] 4) The pedestrian encoding coefficients obtained by training the pedestrian feature information after removing the domain information are combined with the identity information to learn the classifier. At the same time, the consistent matrix and irrelevant matrix of the identity information are introduced to give the classifier a higher discrimination ability;
[0135] 5) Determine the overall objective function of cross-view person re-identification for the identity information matrix;
[0136] 6) Solving the variables to be updated in the overall objective function;
[0137] 7) Design a pedestrian matching solution based on the model that introduces the identity information matrix after removing the domain information.
[0138] The specific implementation process is as follows: first, two dictionaries are learned through low-rank decomposition, one dictionary is used to represent the interference information other than pedestrians in the image, and the other dictionary represents the pedestrian feature information after removing the interference information; then, the pedestrian coding coefficients obtained by training corresponding to the pedestrian feature information after removing the domain information are combined with the identity information to train the classifier. At the same time, in order to enhance the role of identity information, the consistent matrix and irrelevant matrix of identity information are introduced; finally, a pedestrian matching scheme is designed based on the trained classifier.
[0139] Furthermore, the dictionary pair of step 1) includes:
[0140] use represents the training set obtained under camera a, represents the training set obtained under camera b, where m represents the feature dimension and n represents the number of pedestrians. D is the domain information dictionary shared by all cameras, that is, the domain universal dictionary, D t is the pedestrian feature dictionary under a single camera, that is, the domain-invariant dictionary. The domain-general dictionary and the domain-invariant dictionary are integrated together. At this time, the dictionary learning algorithm is shown in formula (1):
[0141]
[0142] In the formula, S a and S b Corresponding training set X on dictionary D a and X b The encoding coefficient matrix, S ta and S tb Corresponding dictionary D t Pedestrian coding coefficients under different viewing angles. The connection between the pedestrian training set of camera a and the domain general dictionary D is established. Same here. and It means to separate the information of domain-general dictionary from the information of domain-invariant dictionary. t ,S a ,S b ,S ta ,S tb ) is the discriminant promotion term of the dictionary Ψ(D,D t ) and the discriminant promotion term Γ(S a ,S b ,S ta ,S tb ). Represents the F-norm.
[0143] Furthermore, the dictionary discrimination promotion items of step 2) include:
[0144] The proposed dictionary discrimination boosting term is:
[0145]
[0146] In the formula, ||D|| * To solve the nuclear norm of the dictionary D, that is, the sum of the singular values of the dictionary D. At the same time, the domain-universal dictionary and the domain-invariant dictionary have different spatial forms. To make the two dictionaries independent of each other, T represents transposition. 1 and α 2 It is a scalar parameter representing the weight information of the contained items.
[0147] Furthermore, the discrimination promotion item of the pedestrian coding coefficient in step 3) includes:
[0148] The proposed discriminant promotion term of the pedestrian coding coefficient is:
[0149] Γ(S a ,S b ,S ta ,S tb )=α 3 {||S a || 2,1 +||S b || 2,1}+α 4 {||S ta || 1 +||S tb || 1} (3)
[0150] In the formula, ||S a || 2,1 and ||S b || 2,1 for l 2,1 norm, used to regularize the matrix, specifically, ||S ta || 1 and ||S tb || 1 for l 1 norm, used to improve the coding coefficient S ta and S tb The sparsity of α 3 and α 4 It is a scalar parameter representing the weight information of the contained items.
[0151] Furthermore, the classifier of step 4) includes:
[0152] Combine the pedestrian encoding coefficients obtained through pedestrian feature information training with identity information to train the classifier W aand W b , realizing the association between the pedestrian appearance features and identity information under two different perspectives a and b. At the same time, the identity consistency matrix and the irrelevant matrix are introduced on the basis of the original model to give the classifier a higher discrimination ability.
[0153] The proposed classifier is:
[0154]
[0155] In the formula, They correspond to the pedestrian identity information matrices under cameras a and b respectively, and k represents the pedestrian category. Represents the identity information matrix between different perspectives. When the visual feature j obtained by camera a and the visual feature k obtained by camera b come from the same person, E(j,k) = 1, otherwise E(j,k) = 0. 5 and α 6 Scalar parameter, representing the weight information of the contained items. The understanding of is: if the training picture is a picture of the same pedestrian under the perspective a, the value obtained is 1; if the training picture is a picture of different pedestrians, the value obtained is 0. Similarly, It represents the b perspective, which has the same function as the a perspective and will not be described here. It represents different viewing angles. If the pedestrian is the same in the two viewing angles, the value is 1. If the pedestrian is different in the two viewing angles, the value is 0.
[0156] Furthermore, the overall objective function of step 5) includes:
[0157] The overall model framework of cross-view person re-identification for identity information matrix is defined as:
[0158]
[0159] In the formula, α 7 and α 8 It is a scalar parameter representing the weight information of the contained items. and Used to control model complexity and avoid overfitting. ||(W a S a ) T (W b S b )|| 1 Yes 1 The constraint of the norm can make (W a S a ) T (W b S b) into a sparse matrix, which further optimizes the computer's matrix calculations and improves the ability to distinguish identity information.
[0160] Furthermore, the variable solving in step 6) includes:
[0161] The variables D, D in formula (5) t ,S a ,S b ,W a ,W b ,S ta ,S tb It cannot be solved directly. Use the iterative shrinkage method to solve different variables separately. The solution of each variable is as follows:
[0162] Fixed D,D t ,S a ,S b ,W b ,S ta ,S tb Update variable W a , then the equation (5) about W a In the objective function, W a The terms can be expressed as:
[0163]
[0164] Formula (6) requires the introduction of an intermediate variable A. The algorithm is as follows:
[0165]
[0166] In formula (7), A is a typical l 1 The norm minimization problem can be solved by using the iterative shrinkage algorithm. After determining A, it is still necessary to introduce an intermediate variable M = W a S ta , then you can a Update. The algorithm is as follows:
[0167]
[0168] The variable M in formula (8) can be directly derived to obtain its solution:
[0169]
[0170] In formula (9), The identity matrix. After the variable M is calculated, W a It can also be directly derived to obtain its solution:
[0171]
[0172] in, is the identity matrix, and d is the size of the dictionary D.
[0173] Using the same method, fix D,D t ,S a ,S b ,W a ,S ta ,S tb Update variable W b , then the equation (5) about W b The objective function can be expressed as:
[0174]
[0175] According to the method in formula (7), an intermediate variable B is introduced to solve the typical l 1 Norm minimization problem. And according to the method of formula (8), an intermediate variable N = W is introduced b S tb After determining B and N, we can calculate W b Direct derivation to obtain its analytical source is:
[0176]
[0177]
[0178] in, is the identity matrix, is the identity matrix, d t It is a dictionary D t size.
[0179] Fixed D,D t ,S b ,W a ,W b ,S ta ,S tb No change, update S a Variable, formula (5) about S a The objective function can be expressed as:
[0180]
[0181] In formula (14), ||S a || 2,1 for l 2,1 Paradigm, find S a The parsed source is:
[0182]
[0183] Among them, Λ 1for The diagonal sparse matrix, S a j Indicates S a The jth column of .
[0184] Use the same method to update the variable S b , fixed D,D t ,S a ,W a ,W b ,S ta ,S tb unchanged, the formula (5) about S b The objective function can be expressed as:
[0185]
[0186] Solve l according to the method in formula (14) 2,1 Paradigm problem, that is, S b Direct derivation yields the following solution:
[0187]
[0188] Among them, Λ 2 for The diagonal sparse matrix, S b j Indicates S b The jth column of .
[0189] Fixed D,D t ,S a ,S b ,W a ,W b ,S tb No change, update S ta Variable, formula (5) about S ta The objective function can be expressed as:
[0190]
[0191] Two intermediate variables C and G need to be introduced in formula (18) to solve l 1 Paradigm problem, equation (18) can be changed to:
[0192]
[0193] S in formula (19) ta It is still not directly solvable, and an intermediate variable O=W needs to be introduced a S ta , rewritten as:
[0194]
[0195] The analytical source of O in formula (20) can be directly obtained as:
[0196]
[0197] in, is the identity matrix.
[0198] After obtaining O, you can ask for S ta , containing S ta The items are:
[0199]
[0200] Therefore, S can be obtained by equation (22): ta Direct derivation yields:
[0201]
[0202] in, is the identity matrix.
[0203] Use the same method to update the variable S tb , fixed D,D t ,S a ,S b ,W a ,W b ,S ta unchanged, the formula (5) about S tb The objective function can be expressed as:
[0204]
[0205] According to the solution method of formula (18), the l in formula (24) can be solved 1 Paradigm problem, equation (24) can be written as:
[0206]
[0207] Formula (25) can be used to calculate S tb Directly solve and obtain its analytical source:
[0208]
[0209] in, is the identity matrix.
[0210] Fixed D t ,S a ,S b ,W a ,W b ,S ta ,S tbKeep the variable D unchanged, and update the variable D. The objective function of formula (5) with respect to D can be expressed as:
[0211]
[0212] Among them, ||D|| * is the nuclear norm, and the singular value threshold algorithm can be used to solve the nuclear norm problem. Introducing an intermediate variable Q = D, equation (27) can be rewritten as:
[0213]
[0214] The analytical source of Q can be found as:
[0215] Q=(α 2 D t D t T +I 7 )\D (29)
[0216] in, is the identity matrix.
[0217] The dictionary D can be calculated by the Lagrange dual method, and its analytical source is:
[0218]
[0219] in, is the identity matrix, Λ 3 is the diagonal matrix obtained from the Lagrange dual variables.
[0220] Finally, according to formula (27), D t Optimize and fix D,S a ,S b ,W a ,W b ,S ta ,S tb constant:
[0221]
[0222] Similarly, D in formula (31) t It can also be obtained by the Lagrange dual method, which is consistent with the solution method of D and will not be repeated here.
[0223] Furthermore, the pedestrian matching scheme of step 7) includes:
[0224] In the test, we use the D,D obtained during training t ,W a ,W b , the following method is used to perform S a ,S b ,Sta ,S tb Solve and match:
[0225]
[0226] In formula (32), X a is the test sample set, S a and S b is the coding coefficient matrix under perspective a and b, S ta and S tb is the encoding coefficient matrix of a specific pedestrian under view a and b. First, fix D,D t ,S ta ,S tb , for S a and S b Update, about S a and S b The terms of the objective function are:
[0227]
[0228]
[0229] S a and S b The analytical source of can be obtained by direct derivation:
[0230]
[0231]
[0232] In formulas (35) and (36), Λ 3 and Λ 4 They are and The diagonal sparse matrix formed by ||S a j || indicates S a The jth column of b j || indicates S b The jth column of S a and S b Then, the same method can be used to find S ta and S tb , whose parsing source is:
[0233]
[0234]
[0235] In formulas (37) and (38), is the identity matrix, A and B are the matrix to solve l 1 The intermediate variable of the paradigm.
[0236] The S obtained by calculating the above test sample set ta and S tb , and the classifier W obtained in the training sample set calculation a and W b To perform distance calculation:
[0237] B a =S ta W a , B b =S tb W b (39)
[0238] Using B in formula (39) a and B b Perform a matching test, assuming B a Set as the array to be tested, B b Set as the target array and calculate B by Euclidean distance a (:,i) and B b (:,j), and the obtained difference values are arranged in ascending order. The obtained order is the Rank matching rate. a (:,i) and B b The difference between (:,j) indicates B a The i-th column of B b The Euclidean distance is calculated for each column of , and the average of these gaps is calculated. The average obtained is the final matching rate.
[0239] The present invention will be further described below in conjunction with specific experimental data.
[0240] In the experiment, the dataset is divided into two different parts, namely the training set under camera a and camera b and the test set under camera a and camera b. The cumulative matching feature (CMC) curve is used to quantitatively evaluate the recognition performance. There are ten parameters in the model, including the dictionary D and D t The size of d and d t and eight scalar parameters, namely α 1 ,α 2 ,α 3 ,α 4 ,α 5 ,α 6 ,α 7 ,α 8 In the whole experiment, the values of the above parameters are set to d = 50, d t =90,α 1 =67,α 2=1.8,α 3 =30,α 4 =2.2,α 5 =1,α 6 =1,α 7 =0.1,α 8 =0.005. The influence of all parameters on the recognition performance is Figure 3-Figure 12 Table 1 shows the performance comparison of the present invention with newer results based on the PRID2011 dataset, with the maximum value bolded.
[0241]
[0242] Table 1: Performance comparison based on newer results on the PRID2011 dataset
[0243] The comparison results show that the method of the present invention has the highest recognition rate at levels 1, 5, and 10, which is even 5.3%, 4.5%, and 3.3% higher than the suboptimal method, respectively.
[0244] The specific implementation modes of the present invention are described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the above implementation modes, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.
Claims
1. A cross-view person re-identification method based on identity information matrix. Features: The steps include: 1) A dictionary pair is learned through low-rank decomposition. One dictionary is used to represent the interference information other than pedestrians in the image, that is, the domain-general dictionary, and the other dictionary represents the pedestrian feature information after removing the interference information, that is, the domain-invariant dictionary; 2) training the discriminant promotion items of the dictionary to make the two different dictionaries independent of each other; 3) Training the discriminative promotion term of pedestrian encoding coefficients, forcing the encoding coefficients of pedestrian images with the same viewpoint to have strong similarity; 4) The pedestrian coding coefficients obtained by training the pedestrian feature information after removing the domain information are combined with the identity information to learn the classifier. At the same time, the consistent matrix and irrelevant matrix of the identity information are introduced to give the classifier a higher discrimination ability; 5) Determine the overall objective function of cross-view person re-identification for the identity information matrix; 6) Solving the variables to be updated in the overall objective function; 7) Design a pedestrian matching solution based on the model that introduces the identity information matrix after removing the domain information.
2. According to the cross-view pedestrian re-identification method for identity information matrix according to claim 1, Features: The domain-general dictionary and domain-invariant dictionary of step 1) include: use represents the training set obtained under camera a, represents the training set obtained under camera b, where m represents the feature dimension and n represents the number of pedestrians; D is the domain information dictionary shared by all cameras, that is, the domain universal dictionary, D t is the pedestrian feature dictionary under a single camera, that is, the domain-invariant dictionary; the domain-general dictionary and the domain-invariant dictionary are integrated together. At this time, the dictionary learning algorithm is shown in formula (1): In the formula, S a and S b Corresponding training set X on dictionary D a and X b The encoding coefficient matrix, S ta and S tb Corresponding dictionary D t Pedestrian coding coefficients under different viewing angles; The connection between the pedestrian training set of camera a and the domain general dictionary D is established. So too; and Indicates the separation of the domain-general dictionary information from the domain-invariant dictionary information; Θ(D,D t ,S a ,S b ,S ta ,S tb ) is the discriminant promotion term of the dictionary Ψ(D,D t ) and the discriminant promotion term Γ(S a ,S b ,S ta ,S tb ); Represents the F-norm.
3. According to the cross-view pedestrian re-identification method for identity information matrix according to claim 2, Features: The dictionary discrimination promotion items of step 2) include: The proposed dictionary discrimination boosting term is: In the formula, ||D|| * To solve the nuclear norm of the dictionary D, that is, the sum of the singular values of the dictionary D; at the same time, the domain-universal dictionary and the domain-invariant dictionary have different spatial forms, so we introduce To make the two dictionaries independent of each other, T represents transposition; α 1 and α 2 It is a scalar parameter representing the weight information of the contained items.
4. According to the cross-view pedestrian re-identification method for identity information matrix according to claim 2, Features: The discrimination promotion item of the pedestrian coding coefficient in step 3) includes: The proposed discriminant promotion term of the pedestrian coding coefficient is: C(S a ,S b ,S ta ,S tb )=a 3 {||S a || 2,1 +||S b || 2,1 }+a 4 {||S ta || 1 +||S tb || 1 } (3) In the formula, ||S a || 2,1 and ||S b || 2,1 for l 2,1 norm, used to regularize the matrix, specifically, ||S ta || 1 and ||S tb || 1 for l 1 norm, used to improve the coding coefficient S ta and S tb The sparsity of α 3 and α 4 It is a scalar parameter representing the weight information of the contained items.
5. According to the method for cross-view person re-identification based on identity information matrix of claim 4, Features: The classifier of step 4) includes: Combine the pedestrian encoding coefficients obtained through pedestrian feature information training with identity information to train the classifier W a and W b , realizing the association between the pedestrian appearance features and identity information under two different perspectives a and b; at the same time, introducing the identity consistency matrix and the irrelevant matrix based on the original model to give the classifier a higher discrimination ability; The proposed classifier is: In the formula, They correspond to the pedestrian identity information matrices under cameras a and b respectively, and k represents the pedestrian category; Represents the identity information matrix between different perspectives. When the visual feature j obtained by camera a and the visual feature k obtained by camera b are from the same person, E(j,k) = 1, otherwise E(j,k) = 0; α 5 and α 6 Scalar parameter, representing the weight information of the contained items; The understanding of is: if the training picture is a picture of the same pedestrian under the perspective a, the value obtained is 1; if the training picture is a picture of different pedestrians, the value obtained is 0; similarly, It represents the b perspective, which has the same function as the a perspective, so I will not go into details here. It represents different viewing angles. If the pedestrian is the same in the two viewing angles, the value is 1. If the pedestrian is different in the two viewing angles, the value is 0.
6. According to the method for cross-view person re-identification based on identity information matrix of claim 5, Features: The overall objective function of step 5) includes: The overall model framework of cross-view person re-identification for identity information matrix is defined as: In the formula, α 7 and α 8 is a scalar parameter, representing the weight information of the contained items; and Used to control model complexity and avoid overfitting; ||(W a S a ) T (W b S b )|| 1 Yes 1 The constraint of the norm can make (W a S a ) T (W b S b ) into a sparse matrix, which further optimizes the computer's matrix calculations and improves the ability to distinguish identity information.
7. According to claim 6, a method for cross-view person re-identification based on identity information matrix, Features: The variable solving of step 6) includes: The variables D, D in formula (5) t ,S a ,S b ,W a ,W b ,S ta ,S tb It cannot be solved directly. Use the iterative shrinkage method to solve different variables separately. The solution of each variable is as follows: Fixed D,D t ,S a ,S b ,W b ,S ta ,S tb Update variable W a , then the equation (5) about W a In the objective function, W a The terms can be expressed as: Formula (6) requires the introduction of an intermediate variable A. The algorithm is as follows: In formula (7), A is a typical l 1 The norm minimization problem can be solved by using the iterative shrinkage algorithm; after determining A, an intermediate variable M = W is still required a S ta , then you can a Update; the algorithm is as follows: The variable M in formula (8) can be directly derived to obtain its solution: In formula (9), The identity matrix; after the variable M is calculated, W a It can also be directly derived to obtain its solution: in, is the identity matrix, d is the size of the dictionary D; Using the same method, fix D,D t ,S a ,S b ,W a ,S ta ,S tb Update variable W b , then the equation (5) about W b The objective function can be expressed as: According to the method in formula (7), an intermediate variable B is introduced to solve the typical l 1 norm minimization problem, and introduce an intermediate variable N = W according to the method of formula (8) b S tb After determining B and N, we can calculate W b Direct derivation to obtain its analytical source is: in, is the identity matrix, is the identity matrix, d t It is a dictionary D t size; Fixed D,D t ,S b ,W a ,W b ,S ta ,S tb No change, update S a Variable, formula (5) about S a The objective function can be expressed as: In formula (14), ||S a || 2,1 for l 2,1 Paradigm, find S a The parsed source is: Among them, Λ 1 for The diagonal sparse matrix, S a j Indicates S a The jth column of Use the same method to update the variable S b , fixed D,D t ,S a ,W a ,W b ,S ta ,S tb unchanged, the formula (5) about S b The objective function can be expressed as: Solve l according to the method in formula (14) 2,1 Paradigm problem, that is, S b Direct derivation yields the following solution: Among them, Λ 2 for The diagonal sparse matrix, S b j Indicates S b The jth column of Fixed D,D t ,S a ,S b ,W a ,W b ,S tb No change, update S ta Variable, formula (5) about S ta The objective function can be expressed as: Two intermediate variables C and G need to be introduced in formula (18) to solve l 1 Paradigm problem, equation (18) can be changed to: S in formula (19) ta It is still not directly solvable, and an intermediate variable O=W needs to be introduced a S ta , rewritten as: The analytical source of O in formula (20) can be directly obtained as: in, is the identity matrix; After obtaining O, you can ask for S ta , containing S ta The items are: Therefore, S can be obtained by equation (22): ta Direct derivation yields: in, is the identity matrix; Use the same method to update the variable S tb , fixed D,D t ,S a ,S b ,W a ,W b ,S ta unchanged, the formula (5) about S tb The objective function can be expressed as: According to the solution method of formula (18), the l in formula (24) can be solved 1 Paradigm problem, equation (24) can be written as: Formula (25) can be used to calculate S tb Directly solve to obtain its analytical source: in, is the identity matrix; Fixed D t ,S a ,S b ,W a ,W b ,S ta ,S tb Keep the variable D unchanged, and update the variable D. The objective function of formula (5) with respect to D can be expressed as: Among them, ||D|| * is the nuclear norm, and the singular value threshold algorithm can be used to solve the nuclear norm problem; introducing an intermediate variable Q = D, formula (27) can be rewritten as: The analytical source of Q can be found as: Q=(α 2 D t D t T +I 7 )\D (29) in, is the identity matrix; The dictionary D can be calculated by the Lagrange dual method, and its analytical source is: in, is the identity matrix, Λ 3 is the diagonal matrix obtained from the Lagrange dual variables; Finally, according to formula (27), D t Optimize and fix D,S a ,S b ,W a ,W b ,S ta ,S tb constant: Similarly, D in formula (31) t It can also be obtained by the Lagrange dual method, which is consistent with the solution method of D and will not be repeated here.
8. According to the method for cross-view person re-identification based on identity information matrix according to claim 7, Features: The pedestrian matching scheme of step 7) includes: In the test, we use the D,D obtained during training t ,W a ,W b , the following method is used to perform S a ,S b ,S ta ,S tb Solve and match: In formula (32), X a is the test sample set, S a and S b is the coding coefficient matrix under view a and b, S ta and S tb is the encoding coefficient matrix of a specific pedestrian under view a and b; first, fix D,D t ,S ta ,S tb , for S a and S b Update, about S a and S b The terms of the objective function are: S a and S b The analytical source of can be obtained by direct derivation: In formulas (35) and (36), Λ 3 and Λ 4 They are and The diagonal sparse matrix formed by ||S a j || indicates S a The jth column of b j || indicates S b The jth column of S a and S b Then, the same method can be used to find S ta and S tb , whose parsing source is: In formulas (37) and (38), is the identity matrix, A and B are the matrix to solve l 1 The intermediate variable of the paradigm; The S obtained by calculating the above test sample set ta and S tb , and the classifier W obtained in the training sample set calculation a and W b To perform distance calculation: B a =S ta W a ,B b =S tb W b (39) Using B in formula (39) a and B b Perform a matching test, assuming B a Set as the array to be tested, B b Set as the target array and calculate B by Euclidean distance a (:,i) and B b (:,j) and sort the obtained difference values in ascending order. The obtained order is the Rank matching rate. a (:,i) and B b The difference between (:,j) indicates B a The i-th column of B b The Euclidean distance is calculated for each column of , and the average of these gaps is calculated. The average obtained is the final matching rate.