Clothing change pedestrian re-identification method and system based on intra-image and inter-image relationship

By preprocessing and constructing an image-internal/external relationship model, and combining the Transformer module to optimize the loss function, the problem of clothing feature interference in pedestrian re-identification during clothing changes is solved, achieving more accurate pedestrian identification.

CN116311377BActive Publication Date: 2025-12-12HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310324819.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-29
Publication Date
2025-12-12
Estimated Expiration
2043-03-29

AI Technical Summary

Technical Problem

Existing pedestrian re-identification methods are easily affected by apparent features such as clothing color when faced with situations involving clothing changes, and most methods fail to effectively utilize the relationships between images, resulting in poor recognition performance.

Method used

By preprocessing all pedestrian images to be identified to ensure they are wearing the same clothing, and by constructing intra-image relationship mining models and inter-image relationship mining models, global and local features are extracted. Feature fusion is then performed using the Transformer module, and the loss function is optimized to improve recognition accuracy.

Benefits of technology

It effectively eliminates the impact of clothing changes on recognition, extracts more discriminative physiological features, and improves the accuracy and robustness of pedestrian re-identification, especially in the utilization of inter-image relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116311377B_ABST
    Figure CN116311377B_ABST
Patent Text Reader

Abstract

The application provides a clothes-changing pedestrian re-identification method and system based on intra-image and inter-image relationships. The method comprises the following steps: obtaining all to-be-identified pedestrian images and preprocessing to obtain a series of pedestrian images wearing the same clothes; constructing an intra-image relationship mining model and an inter-image relationship mining model; dividing the preprocessed to-be-identified pedestrian images into multiple batches; for the current batch, modeling the intra-image relationship by using the intra-image relationship mining model to obtain N fusion features; for the current batch, constructing an inter-image relationship feature by using the inter-image relationship mining model according to the N fusion features, fusing the inter-image relationship feature with the N fusion features respectively to obtain the final feature of each of the N to-be-identified pedestrian images, and judging whether the pedestrian in each to-be-identified pedestrian image is the target pedestrian according to the final feature; for the next batch, repeating the first two steps until the identification of all batches is completed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pedestrian re-identification, and in particular to a clothing-changing pedestrian re-identification method and system based on intra-image and inter-image relationships. BACKGROUND

[0002] Pedestrian re-identification can be regarded as a pedestrian retrieval problem, which aims to retrieve a specific person from a large number of images taken by different cameras and scenes. It mainly faces many challenges such as low image resolution, changes in view angle, posture, light, occlusion, and clothing change. Most current pedestrian re-identification methods are based on the assumption that the clothing of a pedestrian does not change in a short period of time. To realize the landing of the pedestrian re-identification industry, the research on clothing-changing pedestrian re-identification is inevitable.

[0003] The research on clothing-changing pedestrian re-identification is more in line with the actual situation and has stronger practical significance. For example, a fugitive criminal may deliberately change his clothes to evade tracking, and a lost child or old person may change his clothes over time, such as taking off his coat or removing his hat. For clothing-changing pedestrian re-identification, a person's identity is usually determined by his physiological features such as appearance, body shape, etc., rather than apparent features such as clothes, shoes, hairstyle, etc. The key to solving the problem of clothing-changing pedestrian re-identification is to force the model to learn the physiological features of the pedestrian that are not easy to disguise or change, rather than apparent features such as clothing color.

[0004] To cope with the impact and challenges of clothing change on pedestrian re-identification, some research has provided methods to segment the human contour and extract more discriminative body shape features. Ye et al. use joint point features and model the relationship between these points, while also introducing a shape decomposition module to eliminate clothing, which inputs the difference between the regularized global feature and the relationship feature into a self-attention network to let the network automatically separate the clothing feature and the body shape feature. Jin et al. use a dual-stream architecture, which approximates a continuous gait frame from a single input query image in the gait stream to learn rich gait information, and obtains a feature vector through an off-the-shelf network (such as ResNet50) in the ReID stream. Then, the same person is implemented on the features of the two streams to impose high-level semantic consistency constraints, thereby prompting the ReID stream to learn clothing-independent gait movement features.

[0005] However, the existing pedestrian re-identification method still has the following limitations: (1) Most of the existing methods only focus on extracting global features, local features or contour features from an image, and rarely use the relationship between images. Although some works have proposed to model the relationship between images using conditional random fields, these works only model the relationship between a small number of images during training, and the relationship learning has certain limitations. (2) The existing method inevitably encounters the influence of apparent features such as color when facing the clothes changing problem. (3) Most of the existing methods are based on CNN network, but CNN can only use local dependency, and often suffers from information loss due to the use of down-sampling operation. SUMMARY

[0006] In order to solve at least part of the above problems, the present application provides a clothes changing pedestrian re-identification method and system based on intra-image and inter-image relationship.

[0007] In one aspect, the present application provides a clothes changing pedestrian re-identification method based on intra-image and inter-image relationship, comprising:

[0008] Step 1: obtaining all the to-be-identified pedestrian images, and pre-processing all the to-be-identified pedestrian images so that the pedestrians contained in all the to-be-identified pedestrian images are all wearing the same clothes;

[0009] Step 2: constructing an intra-image relationship mining model and an inter-image relationship mining model;

[0010] Step 3: dividing all the pre-processed to-be-identified pedestrian images into multiple batches, each batch containing N to-be-identified pedestrian images;

[0011] Step 4: for the current batch, using the intra-image relationship mining model to model the intra-image relationship of the N to-be-identified pedestrian images respectively to obtain N fusion features; wherein the process of intra-image relationship modeling of each to-be-identified pedestrian image specifically includes: extracting the global feature and the local feature of the to-be-identified pedestrian image, and taking the combination of the global feature and the local feature as the original intra-image feature; constructing the intra-image relationship feature about the to-be-identified pedestrian image according to the global feature and the local feature, and fusing the intra-image relationship feature with the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image;

[0012] Step 5: For the current batch, based on the N fusion features corresponding to the N pedestrian images to be identified, construct the inter-image relationship features between the N pedestrian images to be identified using the image relationship mining model. Fuse the inter-image relationship features with the N fusion features to obtain the final features of each of the N pedestrian images to be identified. Determine whether the pedestrian in each pedestrian image to be identified is the target pedestrian based on the final features of each of the N pedestrian images to be identified.

[0013] Step 6: For the next batch, repeat steps 4 and 5 until all batches have been identified.

[0014] Furthermore, step 1 specifically includes:

[0015] Step 1.1: Select one image from all the images of pedestrians to be identified as the reference image for the target pedestrian;

[0016] Step 1.2: Use the human body analysis model to perform semantic segmentation on the input pedestrian image to be identified, and obtain the pixels belonging to the pedestrian's body in the pedestrian image; the pedestrian's body includes two body parts: the upper garment and the lower garment;

[0017] Step 1.3: Replace the pixels of each part of the pedestrian's body in the target pedestrian reference image with the pixels of the corresponding body parts in the other pedestrian images to be identified, while leaving the pixels in other parts of the other pedestrian images unchanged.

[0018] Furthermore, the human body analysis model is the SCHP model.

[0019] Furthermore, the image-internal relationship mining model includes a CNN model, a human pose estimation module, and a first Transformer module; correspondingly, step 4 specifically includes:

[0020] A CNN model is used to extract global features from the input pedestrian image to be identified;

[0021] A human pose estimation module is used to extract local key point heatmaps from the input pedestrian image to be identified;

[0022] The result of multiplying the global feature and the local key point heatmap is used as the local feature;

[0023] The first Transformer module is used to construct intra-image relation features about the pedestrian image to be identified based on the global features and the local features;

[0024] The result of multiplying the in-image relation features with the original image features is used as the fusion feature of the pedestrian image to be identified.

[0025] Further, in step 4, before constructing the intra-image relationship features, it further comprises:

[0026] The loss function described by formula (1) is used to optimize the global features and the local features.

[0027]

[0028] Wherein, K represents the number of features contained in the local feature V l , is the confidence of the kth key point, represents the heat map of the kth key point, max represents the maximum value operation, and β K+1 = 1 means the confidence of the global feature v K+1 , represents the classification loss function, represents the triple loss function, is the probability that the local feature v k belongs to the identity true value predicted by the classifier, and α is the boundary, represents the distance between the positive feature pair (v ak , v pk ) from the same pedestrian, represents the distance between the negative feature pair (v ak , v nk ) from different pedestrians.

[0029] Further, the key points include one or more of nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left crotch, right crotch, left knee, right knee, left ankle and right ankle.

[0030] Further, the intra-image relationship features are represented by the aggregation representation vector u i shown in formula (2):

[0031]

[0032] Wherein, s(·) is a Softmax function for converting affinity to weight, is a linear projection function for mapping the vector to the v matrix, A ∈ R (K+1)×(K+1) represents the affinity matrix containing the similarity between any two features v i and v j , i ≠ j and i ∈ [1, 2,..., K, K+1], j ∈ [1, 2,..., K, K+1], when i = K+1 or j = K+1, v i or v j represents the global feature, and vice versa, which represents the local feature.

[0033] The affinity matrix A is calculated according to formula (3):

[0034]

[0035] in Let represent the linear projection functions that map the vector to matrices q and k, respectively, where K(·,·) is the inner product function. q represents the scaling factor. i Indicates v i The corresponding matrix, k j Indicates v j The corresponding matrix, T, represents the transpose.

[0036] Furthermore, the image relationship mining model includes a second Transformer module; correspondingly, step 5 specifically includes:

[0037] The second Transformer module is used to construct inter-image relationship features between the N pedestrian images to be identified based on the N fusion features corresponding to the N pedestrian images to be identified using the inter-image relationship mining model.

[0038] Furthermore, the inter-image relationship features are represented by the aggregated representation vector w as shown in formula (4). d To indicate:

[0039]

[0040] Where s(·) is the Softmax function that converts affinity into weights. It is a linear projection function, B∈R N ×N It indicates that it contains any two fused features f d and f e The affinity matrix of similarity between them, where d≠e and d∈[1,2,...,N] and j∈[1,2,...,N];

[0041] The affinity matrix B is calculated according to formula (5):

[0042]

[0043] in Let represent the linear projection functions that map the vector to matrices q and k, respectively, where K(·,·) is the inner product function. q represents the scaling factor. d f d The corresponding matrix, k e f e The corresponding matrix, T, represents the transpose.

[0044] Furthermore, it also includes: constructing the loss function. To optimize the relationship features between images;

[0045]

[0046]

[0047]

[0048]

[0049] Where λ1, λ2, and λ3 represent weights, and L id L represents Identity Loss. tri Indicates Triplet Loss, L C Indicates Center Loss. It is a fusion feature f r The probability of belonging to the true identity predicted by the classifier, where α is the boundary value. Positive feature pairs (f) representing the same pedestrian ar ,f pr The distance between them Negative feature pairs (f) representing different pedestrians ar ,f nr The distance between them; m represents the total number of fused features. Let y represent the fusion feature. t The feature representation center of each class.

[0050] On the other hand, the present invention provides a pedestrian re-identification system for changing clothes based on intra-image and inter-image relationships, comprising: an image preprocessing unit, an intra-image relationship mining unit, an inter-image relationship mining unit, and a recognition unit;

[0051] The image preprocessing unit is used to acquire all pedestrian images to be identified and preprocess all pedestrian images to be identified so that all pedestrians in all pedestrian images to be identified are wearing the same clothes.

[0052] The image-intra relationship mining unit is configured to construct an image-intra relationship mining model, so as to model the image-intra relationship of each of the N to-be-identified pedestrian images contained in the input batch by using the image-intra relationship mining model to obtain N fusion features; wherein the process of modeling the image-intra relationship of each to-be-identified pedestrian image specifically comprises: extracting global features and local features of the to-be-identified pedestrian image, and taking the combination of the global features and the local features as original image-intra features; constructing image-intra relationship features about the to-be-identified pedestrian image according to the global features and the local features, and fusing the image-intra relationship features with the original image-intra features to obtain the fusion features of the to-be-identified pedestrian image;

[0053] The image-intra relationship mining unit is configured to construct an image-intra relationship mining model, so as to model the image-intra relationship of each of the N to-be-identified pedestrian images contained in the input batch by using the image-intra relationship mining model to obtain N fusion features; wherein the process of modeling the image-intra relationship of each to-be-identified pedestrian image specifically comprises: extracting global features and local features of the to-be-identified pedestrian image, and taking the combination of the global features and the local features as original image-intra features; constructing image-intra relationship features about the to-be-identified pedestrian image according to the global features and the local features, and fusing the image-intra relationship features with the original image-intra features to obtain the fusion features of the to-be-identified pedestrian image;

[0054] The recognition unit is configured to determine whether the pedestrians in each of the N to-be-identified pedestrian images contained in the input batch are target pedestrians according to the final features of the N to-be-identified pedestrian images.

[0055] The present application has the following advantages:

[0056] (1) Unlike the method of separating clothing features and identity features by using a self-attention network in the prior art, the present application does not need to separate the clothing features and identity features of pedestrians, but instead adopts preprocessing to preprocess all original to-be-identified pedestrian images into a series of pedestrian images with the same clothes. Since the preprocessed pedestrian images are all wearing the same clothes, even pedestrians with different identities, the clothing features are no longer used as features for identifying the identities of pedestrians. Therefore, the subsequent model design does not need to pay too much attention to the clothing of pedestrians, thereby avoiding the influence of clothing features on the model and making the model no longer rely on color appearance features to identify the identities of pedestrians, so that the designed model can extract more discriminative shape features.

[0057] (2) Compared with the prior method without additional processing of the original image dataset, the application can discard the greatest influencing factor before performing pedestrian feature extraction by adopting a human parsing model to segment the image to obtain the pixels of the body parts (i.e., the two body parts of the upper and lower clothes) that have the greatest impact on pedestrian identity recognition and performing pixel replacement processing, thereby obtaining a series of pedestrian images with the same clothes and eliminating the influence of clothes changing on pedestrian re-identification. At the same time, through this preprocessing method, the model can also focus more on the physiological features that are not easy to change without expanding the dataset, so that the model extracts more discriminative shape features;

[0058] (3) Unlike the prior method of extracting only local information from an image and comparing or extracting only global information and comparing to identify the identity of the pedestrian, the application proposes to construct intra-image relationship features and inter-image relationship features, and then simultaneously utilize the intra-image and inter-image relationship information to extract more discriminative identity features, fully utilizing the semantic information contained in the image.

[0059] (4) Compared with the prior method of obtaining local features by horizontal segmentation, the application adopts human pose estimation to obtain more accurate local key points, and since there is a relationship between the local key points of adjacent joints under the human topological structure, exploring the relationship of all local key points can extract more relationship information.

[0060] (5) Compared with the method of processing the head block in the prior clothes-changing pedestrian re-identification method, the application adopts human pose estimation to additionally extract five key points of the face, including the nose, left eye, right eye, left ear, and right ear, and can extract more discriminative fine-grained features by exploring the relationship between the five key points.

[0061] (6) Compared with other pedestrian re-identification methods based on CNN network, the application utilizes the strong ability of Transformer to obtain long-distance dependency, and since the introduction of multi-head attention, the model can jointly focus on different representation elements, thereby focusing on different parts of the pedestrian under the condition that all images are wearing the same clothes and extracting discriminative shape features. BRIEF DESCRIPTION OF DRAWINGS

[0062] Figure 1 The flowchart of the clothes-changing pedestrian re-identification method based on intra-image and inter-image relationship provided by the embodiment of the application;

[0063] Figure 2 The structural diagram of the intra-image relationship mining model and the inter-image relationship mining model provided by the embodiment of the application. DETAILED DESCRIPTION

[0064] The technical solutions and advantages of the present application will be described clearly below in conjunction with the drawings of the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0065] Embodiment 1

[0066] As shown in the drawings, the embodiment of the present application provides a clothes-changing pedestrian re-identification method based on intra-image and inter-image relationship, comprising the following steps: Figure 1

[0067] S101: acquiring all to-be-identified pedestrian images, and pre-processing all to-be-identified pedestrian images so that all to-be-identified pedestrian images contain pedestrians wearing the same clothes;

[0068] S102: constructing an intra-image relationship mining model and an inter-image relationship mining model;

[0069] S103: dividing all pre-processed to-be-identified pedestrian images into multiple batches, and each batch containing N to-be-identified pedestrian images;

[0070] S104: for the current batch, using the intra-image relationship mining model to model the intra-image relationship of N to-be-identified pedestrian images respectively to obtain N fusion features; wherein the process of modeling the intra-image relationship of each to-be-identified pedestrian image specifically includes: extracting the global feature and the local feature of the to-be-identified pedestrian image, and taking the combination of the global feature and the local feature as the original intra-image feature; constructing the intra-image relationship feature about the to-be-identified pedestrian image according to the global feature and the local feature, and fusing the intra-image relationship feature and the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image;

[0071] S105: for the current batch, according to N fusion features corresponding to N to-be-identified pedestrian images, using the inter-image relationship mining model to construct the inter-image relationship feature about the N to-be-identified pedestrian images, fusing the inter-image relationship feature and the N fusion features respectively to obtain the final feature of each of the N to-be-identified pedestrian images, and judging whether the pedestrian in each to-be-identified pedestrian image is the target pedestrian according to the final feature of each of the N to-be-identified pedestrian images;

[0072] ​Specifically, the identity of a person is usually determined by physiological characteristics rather than apparent characteristics such as clothes. By utilizing the interaction between images to extract the relationship features between images, the significant features that distinguish one pedestrian image from other pedestrian images can be better discovered.

[0073] S106: Repeat the execution of steps S104 to S105 for the next batch until the identification of all batches is completed.

[0074] Unlike the method of separating clothes features and identity features by using self-attention networks in the prior art, the preprocessing method adopted in the embodiment of the present application does not separate the clothes features and identity features of pedestrians. Instead, all input original pedestrian images to be identified are preprocessed into a series of pedestrian images with the same clothes. Since the preprocessed pedestrian images are all wearing the same clothes, even pedestrians with different identities, the clothes features are no longer used as features to distinguish the identities of pedestrians. Therefore, in the subsequent model design, the clothes of pedestrians do not have to be excessively concerned about, thereby avoiding the influence of clothes features on the model and making the model no longer rely on color appearance features to identify pedestrian identities, so that the designed model can extract more distinguishable shape features.

[0075] Moreover, unlike the way of extracting only local information from an image and comparing it or extracting only global information and comparing it to identify the identity of a pedestrian in the prior art, the embodiment of the present application proposes to construct the intra-image relationship features and the inter-image relationship features, and then utilize the intra-image and inter-image relationship information to extract more distinguishable identity features, thereby fully utilizing the semantic information contained in the image.

[0076] Embodiment 2

[0077] On the basis of the above-mentioned embodiments, an image preprocessing method is provided in the embodiment of the present application, so that all the pedestrians contained in the pedestrian images to be identified are wearing the same clothes. The method specifically includes the following steps:

[0078] S201: Select one image from all the pedestrian images to be identified as a target pedestrian reference image; in this embodiment, the first pedestrian image to be identified inputted is selected as the target pedestrian reference image by default.

[0079] S202: Perform semantic segmentation on the input pedestrian image to be identified by using a human body analysis model to obtain the pixels belonging to the pedestrian body in the pedestrian image to be identified; the pedestrian body includes two body parts, i.e., an upper garment and a lower garment (also referred to as trousers); in this embodiment, the human body analysis model is an SCHP model.

[0080] Specifically, all the pedestrian images to be identified in the current batch are denoted as X = [x1, x2,...., xn], and the target pedestrian reference image is denoted as x0. Nwhere N is batch size, x i represents the i-th pedestrian image to be recognized with the size of CxHxW, where C, H, and W represent the number of channels, height, and width, respectively.

[0081] First, the result of semantic segmentation using the human parsing model is represented as S=[s1, s2,...., s N ], s i represents the semantic segmentation result of the image x i , and its size is 1xHxW. In order to identify the semantic information of the pixels in s i , the pixel values of the background, head, upper body, lower body, arm, and leg positions can be set to 0, 1, 2, 3, 4, and 5, respectively.

[0082] Then, the pixels belonging to the pedestrian body (i.e., the upper body and the lower body) are obtained from the above semantic segmentation result. Each pixel of the image x i can be represented as a vector v j with a length of C, so that the image x i has a total of HxW pixel vectors. The pixel set of each body part in the image x i is:

[0083]

[0084]

[0085] where, represents the pixel set of the upper body part, represents the pixel set of the lower body part, U1 is the total number of pixel vectors of the upper body in the image x i , and U2 is the total number of pixel vectors of the lower body in the image x i , represents the j1-th pixel vector, v j2 represents the j2-th pixel vector, 2 represents the index of the upper body, and 3 represents the index of the lower body.

[0086] It can be understood that the values of U1 and U2 can be different in each x i .

[0087] S203: When the remaining pedestrian images to be recognized are input, the pixels of each part of the pedestrian body corresponding to the target pedestrian reference image are respectively replaced into the positions of the pixels of the corresponding body parts of the remaining pedestrian images to be recognized, and the pixels of other positions of the remaining pedestrian images to be recognized remain unchanged.

[0088] Specifically, the pixel set of each body part corresponding to the target pedestrian reference image is stored separately and denoted as G, the pixel set of the body parts (i.e., the upper garment and the lower garment) contained in G is respectively denoted as G upper and G pants ; it is assumed that M is the total number of pixels, for N pedestrian images to be identified, there are M = N x H x W, all the pixel vectors in X are denoted as V X :

[0089]

[0090] wherein, the pixel vector belonging to the upper garment, the pixel vector belonging to the trousers.

[0091] Next, the pixel vector set of the pedestrian images to be identified other than the image G in V X is replaced by G and G respectively, G upper and G pants ; the changed pixel vector can be denoted as V X ':

[0092]

[0093] The other steps are the same as those in Embodiment 1, which will not be described here.

[0094] In the embodiments of the present application, the pixels of the body parts (i.e., the two body parts of the upper garment and the lower garment) that have the greatest impact on pedestrian identity recognition are obtained by using the human body analysis model to segment the image, and the pixel replacement processing is performed, so that the greatest impact factor can be discarded before the pedestrian feature extraction, a series of pedestrian images with the same clothes are obtained, and the influence of clothes changing on pedestrian re-identification is eliminated. At the same time, through this preprocessing method, the model can also be more focused on the physiological features that are not easy to change without expanding the data set, so that the model can extract more discriminative shape features.

[0095] Embodiment 3

[0096] Based on the above embodiments, as shown in Figure 2 , the embodiments of the present application provide a network architecture of an image intra-relation mining model and a network architecture of an image inter-relation mining model, based on the two relation mining models provided in the embodiments, the image intra-relation features and the image inter-relation features can be better constructed. Figure 2 In the embodiments, IP represents an image preprocessing process, INS represents an image intra-relation mining model, and ITS represents an image inter-relation mining model, which are specifically as follows:

[0097] The intra-image relationship mining model comprises a CNN model, a human pose estimation module and a first Transformer module; correspondingly, step S104 specifically comprises: adopting the CNN model to extract a feature map f cnn of the input to-be-identified pedestrian image; kp adopting the human pose estimation module to extract a local key point heat map m g of the input to-be-identified pedestrian image; g obtaining a global feature V cnn through an average pooling operation (g(·)) on the feature map (for convenience of description, denoted as V l = g(f K+1 )); adopting the first Transformer module to construct an intra-image relationship feature of the to-be-identified pedestrian image according to the global feature and the local feature; and multiplying the intra-image relationship feature and the original intra-image feature to obtain a fusion feature of the to-be-identified pedestrian image.

[0098] Specifically, the key points in the local key point heat map comprise one or more of a nose, a left eye, a right eye, a left ear, a right ear, a left shoulder, a right shoulder, a left elbow, a right elbow, a left wrist, a right wrist, a left crotch, a right crotch, a left knee, a right knee, a left ankle and a right ankle; it can be understood that one key point corresponds to one local feature, that is, there can be more than one local feature.

[0099] Further, in order to optimize the global feature and the local feature, before constructing the intra-image relationship feature, there further comprises: optimizing the global feature and the local feature by adopting the loss function in formula (1).

[0100]

[0101] wherein K represents the number of features contained in the local feature V l , is the confidence of the kth key point, represents the heat map of the kth key point, max represents a maximum value operation, and β K+1 = 1 means the confidence of the global feature v K+1 , represents a classification loss function, represents a triple loss function, is the probability that the local feature v k belongs to the true value of the identity predicted by the classifier, and α is a boundary, represents the distance between positive feature pairs (v ak , v pk ) from the same pedestrian, Negative feature pairs (v) representing different pedestrians ak ,v nk The distance between them. It should be noted that classifiers for different local features are not shared.

[0102] based on Figure 2 The image-internal relationship mining model shown uses global features V g and a set of local features V l Simultaneously, the data is input into the first Transformer module for image intra-relation modeling, and the image intra-relation features are represented by the aggregated vector u shown in formula (2). i To indicate:

[0103]

[0104] Where s(·) is the Softmax function that converts affinity into weights. It is a linear projection function that maps a vector to a matrix v, A∈R (K+1)×(K+1) It means that it contains any two features v i and v j The affinity matrix of similarity between them, i≠j and i∈[1,2,...,K,K+1], j∈[1,2,...,K,K+1], when i=K+1 or j=K+1, v i Or v j A single line represents a global feature, while a double line represents a local feature.

[0105] The affinity matrix A is calculated according to formula (3):

[0106]

[0107] in Let represent the linear projection functions that map the vector to matrices q and k, respectively, where K(·,·) is the inner product function. q represents the scaling factor. i Indicates v i The corresponding matrix, k j Indicates v j The corresponding matrix, T, represents the transpose.

[0108] Finally, the obtained image intra-image relation features u i With all corresponding original features v i The features are multiplied separately and then concatenated into a single feature, which is the fused feature of the image. This fused feature makes full use of the relationship information between local key points in the image and is more robust.

[0109] The image relationship mining model includes a second Transformer module. Correspondingly, step S105 specifically includes: using the second Transformer module to construct image relationship features between the N pedestrian images to be identified based on the N fusion features corresponding to the N pedestrian images to be identified, and utilizing the image relationship mining model. For ease of description, the N fusion features corresponding to the N pedestrian images to be identified are denoted as F = [f1, f2, ..., f...]. N ].

[0110] based on Figure 2 The image relationship mining model shown inputs N fused features F into the second Transformer module to model the image relationships. The image relationship features are represented by the aggregated representation vector w shown in formula (4). d To indicate:

[0111]

[0112] Where s(·) is the Softmax function that converts affinity into weights. It is a linear projection function that maps a vector to a matrix v; B∈R N×N It indicates that it contains any two fused features f d and f e The affinity matrix of similarity between them, where d≠e and d∈[1,2,...,N] and j∈[1,2,...,N].

[0113] The affinity matrix B is calculated according to formula (5):

[0114]

[0115] in, Let represent the linear projection functions that map the vector to matrices q and k, respectively, where K(·,·) is the inner product function. q represents the scaling factor. d f d The corresponding matrix, k e f e The corresponding matrix, T, represents the transpose.

[0116] Finally, the obtained relational features w d With all corresponding original features f d Multiply them separately, and use the resulting features as the final features of the corresponding pedestrian image to be identified.

[0117] Furthermore, to improve the discriminative power of deep learning features and minimize intra-class variations while maintaining the separability of features from different classes, this embodiment also constructs a loss function. to optimize the inter-image relationship features. Specifically, as follows:

[0118]

[0119]

[0120]

[0121]

[0122] where λ1, λ2 and λ3 represent weights, L id represents Identity Loss, L tri represents Triplet Loss, L C represents Center Loss, is a fusion feature f r , which belongs to the probability of the identity truth value predicted by the classifier, and α is a boundary, represents the distance between positive feature pairs (f ar , f pr ) from the same pedestrian, represents the distance between negative feature pairs (f ar , f nr ) from different pedestrians; one fusion feature is a sample, in order to minimize the distance between each sample and the corresponding class center in the min-batch, we provide a class center for each class, and m represents the total number of fusion features (i.e. the number of samples), is the feature representation center of the y t th class of fusion features. In this embodiment, λ1, λ2 and λ3 take values of 1, 1 and 0.0005, respectively.

[0123] It is worth noting that when the derivative of the class center is taken, only the pictures of a certain class in the current batch are used to obtain the update amount of the class center. The update strategy of the Center Loss is as shown in formula (10):

[0124]

[0125]

[0126] where δ(condition) takes the following values: when condition is true, it takes a value of 1, otherwise it takes a value of 0.

[0127] Compared with the existing method of obtaining local features by horizontal segmentation, the human pose estimation of the embodiment can obtain more accurate local key points, and since the local key points of adjacent joints are related under the human topological structure, the relationship exploration of all local key points can extract more relationship information. At the same time, compared with the method of processing the head block specially in the existing clothes-changing pedestrian re-identification method, the human pose estimation of the embodiment can additionally extract five key points of the face, including the nose, left eye, right eye, left ear and right ear, and can extract more discriminative fine-grained features by exploring the relationship between the five key points.

[0128] At the same time, compared with other pedestrian re-identification methods based on CNN network, the embodiment utilizes the strong ability of Transformer to obtain long-distance dependency, and since the introduction of multi-head attention, the model can jointly focus on different representation elements, so that different parts of the pedestrian can be focused on in the case that all images are wearing the same clothes, and shape features with discriminative ability can be extracted.

[0129] Embodiment 4

[0130] Corresponding to the above method, the embodiment of the application provides a clothes-changing pedestrian re-identification system based on intra-image and inter-image relationship, which comprises an image preprocessing unit, an intra-image relationship mining unit, an inter-image relationship mining unit and a recognition unit.

[0131] The image preprocessing unit is configured to obtain all the to-be-identified pedestrian images, and pre-process all the to-be-identified pedestrian images so that the pedestrians included in all the to-be-identified pedestrian images are all wearing the same clothes.

[0132] The intra-image relationship mining unit is configured to construct an intra-image relationship mining model, so as to model the intra-image relationship of N to-be-identified pedestrian images contained in an input batch by using the intra-image relationship mining model to obtain N fusion features; wherein the process of modeling the intra-image relationship of each to-be-identified pedestrian image specifically comprises: extracting the global feature and the local feature of the to-be-identified pedestrian image, and taking the combination of the global feature and the local feature as the original intra-image feature; constructing the intra-image relationship feature of the to-be-identified pedestrian image according to the global feature and the local feature, and fusing the intra-image relationship feature and the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image.

[0133] The inter-image relationship mining unit is configured to construct an inter-image relationship mining model, so as to construct inter-image relationship features between N to-be-recognized pedestrian images according to N fusion features corresponding to the N to-be-recognized pedestrian images contained in the input batch by using the inter-image relationship mining model, and fuse the inter-image relationship features with the N fusion features respectively to obtain final features of the N to-be-recognized pedestrian images respectively.

[0134] The recognition unit is configured to determine whether a pedestrian in each to-be-recognized pedestrian image is a target pedestrian according to the final feature of each to-be-recognized pedestrian image contained in the input batch.

[0135] It should be noted that the clothing-changing pedestrian re-identification system based on intra-image and inter-image relationships provided by the embodiments of the present application is to realize the above-mentioned method embodiments, and the functions thereof can be referred to the above-mentioned method embodiments, which will not be described here.

[0136] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for pedestrian re-identification based on intra-image and inter-image relations, characterized in that, The method comprises the following steps: Step 1: obtaining all to-be-identified pedestrian images, and pre-processing all to-be-identified pedestrian images so that all to-be-identified pedestrian images contain pedestrians wearing the same clothes; Step 2: constructing an intra-image relationship mining model and an inter-image relationship mining model; Step 3: dividing all pre-processed to-be-identified pedestrian images into multiple batches, each batch containing N to-be-identified pedestrian images; Step 4: for the current batch, using the intra-image relationship mining model to model the intra-image relationship of the N to-be-identified pedestrian images respectively to obtain N fusion features; wherein the intra-image relationship modeling process of each to-be-identified pedestrian image specifically comprises: extracting the global feature and the local feature of the to-be-identified pedestrian image, and taking the combination of the global feature and the local feature as the original intra-image feature; constructing the intra-image relationship feature of the to-be-identified pedestrian image according to the global feature and the local feature, and fusing the intra-image relationship feature and the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image; Step 5: for the current batch, according to the N fusion features corresponding to the N to-be-identified pedestrian images, using the inter-image relationship mining model to construct the inter-image relationship feature of the N to-be-identified pedestrian images, fusing the inter-image relationship feature and the N fusion features respectively to obtain the final feature of the N to-be-identified pedestrian images respectively, and judging whether the pedestrian in each to-be-identified pedestrian image is the target pedestrian according to the final feature of the N to-be-identified pedestrian images respectively; Step 6: for the next batch, repeating steps 4 to 5 until the identification of all batches is completed.

2. The image-based inter-and-intra relation based pedestrian re-identification method of changing clothes according to claim 1, characterized in that, Step 1 specifically comprises: Step 1.1: selecting one image from all to-be-identified pedestrian images as a target pedestrian reference image; Step 1.2: using a human body analysis model to perform semantic segmentation on the input to-be-identified pedestrian image to obtain pixels belonging to the pedestrian body in the to-be-identified pedestrian image; the pedestrian body includes two body parts of a top and a bottom; Step 1.3: replacing the pixels of each part of the pedestrian body corresponding to the target pedestrian reference image to the positions of the pixels of the corresponding body parts of the remaining to-be-identified pedestrian images, and keeping the pixels of other positions of the remaining to-be-identified pedestrian images unchanged. 3.The image-based cross-dressing person re-identification method according to claim 2, wherein, The human body analysis model is an SCHP model.

4. The image-in and image-out relationship based clothes changing pedestrian re-identification method according to claim 1, characterized in that, The intra-image relationship mining model comprises a CNN model, a human body pose estimation module and a first Transformer module; Correspondingly, step 4 specifically comprises: using the CNN model to extract the global feature of the input to-be-identified pedestrian image; using the human body pose estimation module to extract the local key point heat map of the input to-be-identified pedestrian image; multiplying the global feature and the local key point heat map to obtain the local feature; using the first Transformer module to construct the intra-image relationship feature of the to-be-identified pedestrian image according to the global feature and the local feature; multiplying the intra-image relationship feature and the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image.

5. The image-in and image-out relationship based pedestrian re-identification method of changing clothes according to claim 4, characterized in that, In step 4, before constructing the intra-image relationship feature, it further comprises: The loss function described by formula (1) is used to optimize the global feature and the local feature; where K denotes the number of local features V l contained in V β k = max(m kp [k]) ∈ [0, 1] is the confidence of the k-th keypoint, m kp [k] denotes the heat map of the k-th keypoint, max denotes the max operation, β K+1 = 1 refers to the confidence of the global feature v K+1 , denotes the classification loss function, denotes the triplet loss function, is the probability that the local feature v k belongs to the ground truth identity predicted by the classifier, and a is the margin, represents the distance between positive feature pairs (v ak , v pk ) from the same pedestrian, represents the distance between negative feature pairs (v ak , v nk ) from different pedestrians. 6.The image-based cross-dressing person re-identification method according to the relationship between intra-image and inter-image relationship of claim 4, characterized in that, The intra-image relational feature is represented by an aggregated representation vector u shown in equation (2) i is represented as: where s(·) is a Softmax function that transforms affinities into weights, is a linear projection function that maps vectors to the matrix v, A e R (K+1)×(K+1) denotes an affinity matrix that contains the similarity between any two features v i and v j , i≠j and i e [1, 2,..., K, K+1], j e [1, 2,..., K, K+1], when i = K+1 or j = K+1, v i or v j denotes a global feature, otherwise, it denotes a local feature; The affinity matrix A is calculated according to formula (3): wherein, denote linear projection functions that map a vector to the q-matrix, k-matrix, respectively, K(·, ·) is an inner product function, denotes a scaling factor, q i denotes v i the corresponding matrix, k j denotes v j the corresponding matrix, T denotes the transpose.

7. The image-in and image-out relationship based clothes changing pedestrian re-identification method according to claim 1, characterized in that, The inter-image relationship mining model comprises a second Transformer module; Correspondingly, step 5 specifically comprises: The second Transformer module is used to construct the inter-image relationship feature between the N to-be-identified pedestrian images according to the N fusion features corresponding to the N to-be-identified pedestrian images by using the inter-image relationship mining model.

8. The image-in and image-out relationship based pedestrian re-identification method of changing clothes according to claim 7, characterized in that, The inter-image relationship feature is represented by an aggregated representation vector w shown in equation (4) d is represented as: where s(·) is a Softmax function that transforms affinities into weights, is a linear projection function that maps a vector to a matrix v, B e R N×N denotes an affinity matrix that contains the similarity between any two fused features f d and f e , d≠e and d e [1, 2,..., N] and j e [1, 2,..., N]. The affinity matrix B is calculated according to formula (5): wherein, denote linear projection functions that map a vector to the q-matrix, k-matrix, respectively, K(·, ·) is an inner product function, denotes a scaling factor, q d denotes f d the corresponding matrix, k e denotes f e the corresponding matrix, T denotes the transpose.

9. The image-in and image-out relationship based clothes changing pedestrian re-identification method according to claim 8, characterized in that, Further comprising: Constructing loss function to optimize inter-image relationship features; where λ1, λ2 and λ3 represent weights, L id represents Identity Loss, L tri represents Triplet Loss, L C represents Center Loss, p fr is the fused feature f r is the probability of the ground truth identity predicted by the classifier, and a is the margin, represents the distance between positive feature pairs (f ar , f pr ) from the same pedestrian, represents the distance between negative feature pairs (f ar , f nr ) from different pedestrians; and m represents the total number of fused features, represents the feature representation center of the y t th class of fused features.

10. A system for person re-identification of a person changing clothes based on intra-image and inter-image relationships, the system comprising: Including: An image preprocessing unit, an intra-image relationship mining unit, an inter-image relationship mining unit and an identification unit; The image preprocessing unit is configured to acquire all to-be-identified pedestrian images, and pre-process all to-be-identified pedestrian images so that all to-be-identified pedestrian images contain pedestrians wearing the same clothes; The intra-image relationship mining unit is configured to construct an intra-image relationship mining model, so as to use the intra-image relationship mining model to model the intra-image relationship of N to-be-identified pedestrian images contained in an input batch to obtain N fusion features; wherein the process of modeling the intra-image relationship of each to-be-identified pedestrian image specifically comprises: extracting the global feature and the local feature of the to-be-identified pedestrian image, and taking the combination of the global feature and the local feature as an original intra-image feature; constructing the intra-image relationship feature of the to-be-identified pedestrian image according to the global feature and the local feature, and fusing the intra-image relationship feature and the original intra-image feature to obtain the fusion feature of the to-be-identified pedestrian image; The inter-image relationship mining unit is configured to construct an inter-image relationship mining model, so as to use the inter-image relationship mining model to construct the inter-image relationship feature between N to-be-identified pedestrian images according to N fusion features corresponding to the N to-be-identified pedestrian images contained in an input batch, and fuse the inter-image relationship feature and the N fusion features to obtain the final feature of the N to-be-identified pedestrian images respectively; The identification unit is configured to determine whether the pedestrian in each to-be-identified pedestrian image is a target pedestrian according to the final feature of the N to-be-identified pedestrian images contained in an input batch.

Citation Information

Patent Citations

  • Dressing pedestrian re-identification and retrieval method based on multi-mode intelligent perception and fusion

    CN114998934A

  • Multi-level relation analysis and mining method for image-text data

    CN115098646A