Pedestrian re-identification method and system based on data enhancement

By constructing the expanded data set and training the CAL model, the problem of low generalization ability of pedestrian re-identification model in the existing technology in the cross-garment change scenarios is solved, which improves recognition accuracy and reduces misidentification.

CN120014668APending Publication Date: 2025-05-16BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510011199.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing data-enhanced pedestrian re-identification technology has low generalization ability in cross-garment changes scenarios, resulting in a decrease in the recognition accuracy of the model and easily lead to misidentification of pedestrians wearing similar clothes.

Method used

By constructing the expanded data set, including the original re-identification data set of pedestrians with similar clothing pedestrians, training the CAL model to enhance the generalization ability of the re-identification model in cross-garment change scenarios.

Benefits of technology

The recognition accuracy of the pedestrian re-identification model is improved, the model's generalization ability in cross-clothing changes scenarios is enhanced, and the misidentification of pedestrians wearing similar clothes is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014668A_ABST
    Figure CN120014668A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a pedestrian re-identification method and system based on data enhancement. The method comprises the steps of obtaining a to-be-recognized image, wherein the to-be-recognized image comprises a target pedestrian; based on an image database and the to-be-recognized image, a re-recognition model is adopted to re-recognize the target pedestrian, a recognition result is obtained, and the recognition result comprises at least one target image which is obtained from the image database and comprises the target pedestrian, the re-identification model is a model which is obtained by training a CAL model by adopting an expanded data set and is used for re-identifying pedestrians, and the expanded data set is composed of an original clothes changing pedestrian re-identification data set and a similar clothes pedestrian data set; the similar clothes pedestrian data set comprises images obtained after pedestrians change similar clothes in a plurality of images in the original clothes changing pedestrian re-identification data set. The method is used for improving the recognition accuracy of the re-recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision, and in particular to a method and system for pedestrian re-identification based on data enhancement. Background Art

[0002] Traditional pedestrian re-identification technology mainly relies on the appearance features of pedestrians (such as clothing color, style, texture, etc.) to realize the recognition and tracking of pedestrians. However, once pedestrians change their clothes, their appearance features change dramatically, and the accuracy of pedestrian recognition methods based on these features will drop significantly. Therefore, Cloth-Changing Person Re-Identification (CC-ReID) technology came into being.

[0003] In the existing technology, CC-ReID technology often expands the number of data sets used to train pedestrian re-identification models through data enhancement, so that the trained pedestrian re-identification model can focus on extracting identity features that are not related to the pedestrian's clothing (such as the pedestrian's posture, contour, etc.) to achieve accurate recognition of pedestrians.

[0004] However, since the clothing of different pedestrians is obviously different in the existing datasets generated based on the idea of ​​data enhancement, the re-identification model will still focus on clothing features during the training process. There is a problem of unbalanced feature expression of the trained re-identification model, which leads to low generalization ability of the model and easily causes misidentification of pedestrians wearing similar clothes. Summary of the invention

[0005] The embodiments of the present application provide a pedestrian re-identification method and system based on data enhancement, which can enhance the generalization ability of the re-identification model in scenes with different clothing changes, thereby achieving the technical effect of improving the recognition accuracy of the model.

[0006] In a first aspect, an embodiment of the present application provides a pedestrian re-identification method based on data enhancement, comprising:

[0007] Acquire an image to be identified, wherein the image to be identified includes a target pedestrian;

[0008] Based on the image database and the image to be identified, re-identify the target pedestrian using a re-identification model to obtain a recognition result, wherein the recognition result includes at least one target image including the target pedestrian obtained from the image database;

[0009] Among them, the re-identification model is a model for re-identifying pedestrians obtained by training the CAL model with an expanded data set. The expanded data set is composed of an original pedestrian re-identification data set and a pedestrian data set with similar clothing. The pedestrian data set with similar clothing includes images obtained after pedestrians in multiple images in the original pedestrian re-identification data set are changed into similar clothing.

[0010] In a possible implementation, the method further includes:

[0011] The clothing of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset is changed into the same clothing, so as to obtain the similar clothing pedestrian dataset composed of multiple images after the clothing-changing; the original clothing-changing pedestrian re-identification dataset includes multiple images containing pedestrians;

[0012] Merging the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset;

[0013] Constructing the CAL model;

[0014] The CAL model is trained according to the expanded data set to obtain the re-identification model.

[0015] In a possible implementation, during the model training process of the CAL model, a joint loss function determined based on the Centroid loss function and the CAL loss function is selected as the loss function of the model.

[0016] In a possible implementation, changing the clothes of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset into the same clothes includes:

[0017] For each of the multiple images in the original clothing-changing pedestrian re-identification dataset, obtain a masked image, a binary mask, and a human posture key point map corresponding to the image;

[0018] Based on the pre-acquired image of the clothing to be changed, acquiring a target clothing image that fits the human body in the image;

[0019] Encoding is performed according to the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing;

[0020] According to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing, a changed-dress image corresponding to the image is obtained.

[0021] In a possible implementation, the step of obtaining a masked image, a binary mask, and a human body posture key point map corresponding to the image includes:

[0022] Processing the image based on the SCHP algorithm to generate a pedestrian segmentation mask;

[0023] According to the pedestrian segmentation mask, acquiring the masked image and the binary mask;

[0024] Based on the OpenPose algorithm, the human body posture key points in the image are extracted to obtain the human body posture key point map.

[0025] In a possible implementation, the step of obtaining a target clothing image that fits the human body in the pre-acquired clothing image includes:

[0026] Based on the pre-acquired image of the clothing to be changed, a TPS algorithm is used to obtain a preliminary deformed image of the clothing to be changed; wherein the image of the clothing to be changed is an image of the clothing in a flat state;

[0027] According to the masked image, the initially deformed clothing image to be changed, and the human body posture key point image, the Unet algorithm is used to obtain the target clothing image that fits the human body in the image.

[0028] In a possible implementation, obtaining the dressed-up image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map, and the encoded target clothing includes:

[0029] Obtaining preset text prompt information, where the text prompt information is used to prompt a pedestrian in the image to change clothes;

[0030] The text prompt information, random noise, the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing are input into the inpainting pipeline of the stable diffusion model for dressing processing to obtain the dressed-up image corresponding to the image.

[0031] In a possible implementation manner, before merging the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset, the method further includes:

[0032] The low-quality images in the similar clothing pedestrian dataset are removed.

[0033] In a second aspect, an embodiment of the present application provides a pedestrian re-identification device based on data enhancement, comprising:

[0034] A first acquisition unit, used to acquire an image to be identified, wherein the image to be identified includes a target pedestrian;

[0035] a recognition unit, configured to re-recognize the target pedestrian using a re-recognition model based on an image database and the image to be recognized, to obtain a recognition result, wherein the recognition result includes at least one target image including the target pedestrian obtained from the image database;

[0036] Among them, the re-identification model is a model for re-identifying pedestrians obtained by training the CAL model with an expanded data set. The expanded data set is composed of an original pedestrian re-identification data set and a pedestrian data set with similar clothing. The pedestrian data set with similar clothing includes images obtained after pedestrians in multiple images in the original pedestrian re-identification data set are changed into similar clothing.

[0037] In a possible implementation, the pedestrian re-identification device based on data enhancement further includes:

[0038] The clothing changing unit is used to change the clothing of pedestrians in multiple images in the original clothing changing pedestrian re-identification data set into the same clothing, so as to obtain the pedestrian data set of similar clothing composed of multiple images after clothing change; the original clothing changing pedestrian re-identification data set includes multiple images containing pedestrians

[0039] An expansion unit, used for merging the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset;

[0040] Construct the unit and build the CAL model;

[0041] A training unit is used to train the CAL model according to the expanded data set to obtain the re-identification model.

[0042] In a possible implementation, in the training unit, during the model training of the CAL model, a joint loss function determined based on the Centroid loss function and the CAL loss function is selected as the loss function of the model.

[0043] In a possible implementation, the clothing changing unit includes:

[0044] A first acquisition module is used to acquire, for each of the multiple images in the original clothing-changing pedestrian re-identification dataset, a masked image, a binary mask and a human posture key point map corresponding to the image;

[0045] A second acquisition module is used to acquire a target clothing image that fits the human body in the image based on the pre-acquired image of the clothing to be changed;

[0046] A third acquisition module is used to encode the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing;

[0047] The fourth acquisition module is used to acquire the changed-dress image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing.

[0048] In a possible implementation manner, the first acquisition module is specifically configured to:

[0049] Processing the image based on the SCHP algorithm to generate a pedestrian segmentation mask;

[0050] According to the pedestrian segmentation mask, acquiring the masked image and the binary mask;

[0051] Based on the OpenPose algorithm, the human body posture key points in the image are extracted to obtain the human body posture key point map.

[0052] In a possible implementation manner, the second acquisition module is specifically configured to:

[0053] Based on the pre-acquired image of the clothing to be changed, a TPS algorithm is used to obtain a preliminary deformed image of the clothing to be changed; wherein the image of the clothing to be changed is an image of the clothing in a flat state;

[0054] According to the masked image, the initially deformed clothing image to be changed, and the human body posture key point image, the Unet algorithm is used to obtain the target clothing image that fits the human body in the image.

[0055] In a possible implementation, the fourth acquisition module is specifically configured to:

[0056] Obtaining preset text prompt information, where the text prompt information is used to prompt a pedestrian in the image to change clothes;

[0057] The text prompt information, random noise, the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing are input into the inpainting pipeline of the stable diffusion model for dressing processing to obtain the dressed-up image corresponding to the image.

[0058] In a possible implementation, the pedestrian re-identification device based on data enhancement further includes:

[0059] The filtering unit is used to remove low-quality images from the similar clothing pedestrian data set.

[0060] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0061] The memory stores computer-executable instructions;

[0062] The processor executes the computer-executable instructions stored in the memory, so that the processor performs the above first aspect and / or various possible pedestrian re-identification methods based on data enhancement in the first aspect.

[0063] In a fourth aspect, an embodiment of the present application provides a pedestrian re-identification system based on data enhancement, including: an electronic device, a terminal device;

[0064] The electronic device is used to execute the above first aspect and / or various possible pedestrian re-identification methods based on data enhancement in the first aspect to obtain a recognition result; and the terminal device is used to output the recognition result.

[0065] The fifth embodiment of the present application provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed by a processor, they are used to implement the first aspect above and / or various possible pedestrian re-identification methods based on data enhancement of the first aspect.

[0066] In a sixth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible pedestrian re-identification methods based on data enhancement in the first aspect.

[0067] The embodiment of the present application provides a pedestrian re-identification method and system based on data enhancement, which obtains an image to be identified including a target pedestrian; based on an image database and the image to be identified, a model for re-identifying pedestrians is obtained by training a CAL model based on a data set after expanding a data set of pedestrians in similar clothing, and re-identifies the target pedestrian, thereby obtaining at least one target image including the target pedestrian obtained from the image database, thereby enhancing the generalization ability of the re-identification model in scenes across clothing changes, thereby achieving the technical effect of improving the recognition accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0069] Figure 1 A flowchart of a pedestrian re-identification method based on data enhancement provided in Example 1 of the present application;

[0070] Figure 2A flowchart of a pedestrian re-identification method based on data enhancement provided in Embodiment 2 of the present application;

[0071] Figure 3 A schematic diagram of replacing similar clothing in the method for generating a changed clothing image provided by the present application;

[0072] Figure 4 A schematic diagram of the principle of a specific method for generating a changed-clothing image provided in this application;

[0073] Figure 5 A schematic diagram of the structure of a pedestrian re-identification system based on data enhancement provided in Embodiment 3 of the present application;

[0074] Figure 6 The system architecture of the specific pedestrian re-identification system based on data enhancement provided in the fourth embodiment of the present application;

[0075] Figure 7 It is a schematic diagram of specific functional modules in the business logic layer of the system architecture;

[0076] Figure 8 A schematic diagram of the structure of a pedestrian re-identification device based on data enhancement provided in Embodiment 5 of the present application;

[0077] Fig. 9 A schematic diagram of the structure of a pedestrian re-identification device based on data enhancement provided in Example 6 of the present application;

[0078] Fig.10 This is a schematic diagram of the structure of an electronic device provided in Example 7 of the present application.

[0079] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0080] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0081] In order to facilitate understanding of the technical solution in this application, the following is a detailed introduction to the background technology:

[0082] In scenes with high crowd density, low camera resolution and limited shooting angle (such as large amusement parks), it is difficult for cameras to clearly capture the frontal image of a person's face, and traditional face recognition technology is difficult to play a role. In this case, pedestrian re-identification technology becomes an effective supplementary means.

[0083] Pedestrian re-identification technology recognizes and tracks targets by analyzing the pedestrian's body shape, contour shape, and other physical characteristics. However, existing re-identification technologies fail to fully consider the possible scenarios of pedestrians changing clothes in reality. This limitation will have a significant impact on the recognition accuracy of the re-identification model. Especially in application scenarios that require continuous identification and tracking of pedestrians across time and space, changing clothes is very common. For example, tourists may change their clothes when entering indoors from outdoors; in long-term data analysis, pedestrians' clothing will change at different times. Therefore, it is particularly important to introduce CC-ReID technology that can handle pedestrians changing clothes.

[0084] In the prior art, the datasets used to train the person re-identification model have the problem of small dataset size and limited coverage of scenarios, making it difficult to simulate the diverse clothing changes and complex human postures in real applications. To address this problem, the prior art often expands the number of datasets used to train the person re-identification model through data enhancement, so that the trained person re-identification model can focus on extracting identity features that are not related to the pedestrian's clothing (such as the pedestrian's posture, outline, etc.) to achieve accurate recognition of pedestrians.

[0085] In the prior art, there are methods for expanding the data set used to train the re-identification model based on the idea of ​​data enhancement. For example, the clothing-changing image generation method based on Generative Adversarial Networks (GAN) mainly obtains multiple images of the image with different clothing by changing the clothing of the pedestrians in each image in the data set to clothing of different styles, so as to expand the number and diversity of the data set. Although this method expands the data set to a certain extent, there are obvious differences in the clothing of different pedestrians in the data set. During the training process, the re-identification model will still focus on clothing features that are irrelevant to the identity of the pedestrians. There is a problem of unbalanced feature expression of the re-identification model after training, which leads to low generalization ability of the model and easy misidentification of pedestrians wearing similar clothes.

[0086] In addition, the images after changing clothes generated by existing methods often have problems such as geometric structure distortion, texture blur or posture abnormality, resulting in the generated low-quality images interfering with model training, and then making the model unable to accurately extract image features when actually re-identifying pedestrians, resulting in low recognition accuracy.

[0087] In view of the technical problems in the above-mentioned background technology, the inventors found that the clothing of pedestrians in the images in the data set is replaced with uniform clothing to construct a similar clothing pedestrian data set in the process of studying the pedestrian re-identification method based on data enhancement. Since the clothing features of the samples in the constructed similar clothing pedestrian data set are more evenly distributed, when training the model based on the similar clothing pedestrian data set, other features that are not related to the pedestrian's clothing can be focused on, which can enhance the generalization ability of the re-identification model obtained through training in scenes with cross-clothing changes, so as to achieve the technical effect of improving the recognition accuracy of the model.

[0088] Based on the technical conception of the above-mentioned inventors, the pedestrian re-identification method based on data enhancement provided by the present application obtains an image to be identified including a target pedestrian; based on an image database and the image to be identified, a model for re-identifying pedestrians is obtained by training a CAL model based on a data set after expanding a data set of pedestrians in similar clothing, and the target pedestrian is re-identified to obtain at least one target image including the target pedestrian obtained from the image database, thereby enhancing the generalization ability of the re-identification model in scenarios across clothing changes, thereby achieving the technical effect of improving the recognition accuracy of the model.

[0089] The pedestrian re-identification method based on data enhancement provided in this application can be applied to application scenarios of pedestrian re-identification such as intelligent search for people and identification of unauthorized persons. This application does not limit the specific application scenarios of this solution.

[0090] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0091] Figure 1 A flowchart of a pedestrian re-identification method based on data enhancement provided in the first embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:

[0092] S101: Acquire an image to be identified, where the image to be identified includes a target pedestrian.

[0093] In this scheme, when re-identifying pedestrians, it is necessary to use a pedestrian re-identification model trained based on a data set after expanding a data set of pedestrians in similar clothing, and obtain at least one target image including the target pedestrian in the image to be identified from the image database.

[0094] In this step, the image to be identified is provided by the user, including the image of the target pedestrian to be identified.

[0095] S102. Based on the image database and the image to be identified, a re-identification model is used to re-identify the target pedestrian to obtain an identification result, wherein the re-identification model is a model for re-identifying pedestrians obtained by training a Clothes-based Adversarial Loss (CAL) model using an expanded data set, and the expanded data set is composed of an original re-identification data set of pedestrians who have changed clothes and a data set of pedestrians in similar clothes.

[0096] In this step, based on multiple images including different pedestrians stored in the image database, a re-recognition model is used to re-recognize the target image included in the image to be recognized, and obtain a recognition result. The images stored in the image database are obtained based on historical monitoring data collected by image acquisition devices such as surveillance cameras; the recognition result includes at least one target image including the target pedestrian obtained from the image database.

[0097] Specifically, the re-identification model extracts features of the target pedestrian in the image to be identified, obtains the target features of the target pedestrian in the image to be identified, extracts features of the pedestrian in each image in the image database, obtains the features of the pedestrian in each image in the image database, and then compares the target features with the features of the pedestrian in each image in the image database, and determines the image in the image database whose features of the pedestrian in the image are similar to the target features as the target image. Based on the obtained at least one target image, a recognition result is obtained.

[0098] It is worth noting that the re-identification model is obtained by training the CAL model based on the expanded dataset, which can more fully extract features related to the identity of pedestrians. Among them, the expanded dataset consists of the original clothing-changing pedestrian re-identification dataset and the similar clothing pedestrian dataset; the similar clothing pedestrian dataset includes images obtained by changing similar clothing from multiple images of pedestrians in the original clothing-changing pedestrian re-identification dataset.

[0099] The data enhancement-based pedestrian re-identification method provided in the embodiment of the present application obtains an image to be identified including a target pedestrian; based on an image database and the image to be identified, a model for re-identifying pedestrians is obtained by training a CAL model based on a data set after expanding a data set of pedestrians in similar clothing, and the target pedestrian is re-identified to obtain at least one target image including the target pedestrian from the image database, thereby enhancing the generalization ability of the re-identification model in scenes across clothing changes, thereby achieving the technical effect of improving the recognition accuracy of the model.

[0100] Furthermore, Figure 2The flowchart of the pedestrian re-identification method based on data enhancement provided in the second embodiment of the present application is as follows. Based on the above embodiment, the pedestrian re-identification method based on data enhancement provided in the present application further includes:

[0101] S201, changing the clothes of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset into the same clothes, and obtaining a similar clothing pedestrian dataset consisting of multiple clothing-changing images.

[0102] In this step, multiple images in the original clothing-changing pedestrian re-identification dataset need to be processed with clothing-changing to construct a dataset of pedestrians in similar clothing.

[0103] The original clothing-changing pedestrian re-identification dataset is at least one open source dataset for clothing-changing re-identification; exemplary, the open source dataset for clothing-changing re-identification may include: Person Reid Clothes-Changing (PRCC) dataset, Long-Term Clothes-Changing (LTCC) dataset, DeepChange dataset, etc. This application does not restrict the selection of the original clothing-changing pedestrian re-identification dataset.

[0104] Specifically, it is necessary to process the images in the dataset to obtain the key information of the pedestrians in the images, and then use the image generation model to change the clothes of the pedestrians in the images based on the key information. Among them, the key information includes the posture of the human body. Optionally, for an original clothing-changing pedestrian re-identification dataset, the pedestrians in some images in the dataset can be changed; or the pedestrians in all images in the dataset can be changed.

[0105] In a specific implementation, the present application also provides a method for generating a clothes-changed image. Figure 3 A schematic diagram of replacing similar clothing in the method for generating a changed clothing image provided by the present application, such as Figure 3 As shown in the figure, the original image is an image in the original clothing-changing pedestrian re-identification dataset. The image after clothing-changing not only retains the identity features of the original image (such as the facial features of the person), but also retains the posture of the human body in the original image, and the changed clothes can also fit the posture of the human body.

[0106] Specifically, the following method can be used to change the clothes of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset into the same clothes:

[0107] S2011. For each of the multiple images in the original clothing-changing pedestrian re-identification dataset, obtain a masked image, a binary mask, and a human posture key point map corresponding to the image.

[0108] In this step, for each image obtained from the original clothing-changing pedestrian re-identification dataset, it is necessary to obtain multiple key information corresponding to the image, including the masked image, the binary mask, and the human posture key point map. Among them, the masked image refers to the image obtained by masking the clothing of the pedestrian in the image; the binary mask is the binary code obtained by encoding the mask with an encoder; the human posture key point map refers to the map that presents the human posture by marking the key points of the human posture.

[0109] In a specific implementation, the above key information may be obtained in the following manner:

[0110] Step 1: Process the image based on the Self-correction-human-parsing (SCHP) algorithm to generate a pedestrian segmentation mask.

[0111] In this step, the semantic segmentation task of pedestrians in the image is carried out based on SCHP. First, a pedestrian segmentation map with 18 category labels is obtained, and then a pedestrian segmentation mask is generated based on the pedestrian segmentation map to indicate the position that needs to be masked in the image. Among them, the category label is used to indicate the type of human body parts of pedestrians, such as tops, pants, skirts, arms, etc.; the position that needs to be masked is the human body part labeled as clothing type in the human body segmentation map.

[0112] It should be understood that the pedestrian segmentation map is a visual presentation of different parts of the human body after division, and the labels of different parts of the human body are represented in the form of different colors. For example, in the pedestrian segmentation map, the top is represented by red; the skirt is represented by green, and so on.

[0113] Step 2: Obtain the masked image and the binary mask based on the pedestrian segmentation mask.

[0114] In this step, the pedestrian segmentation mask is used to mask the image to obtain a masked image that masks the clothing parts of the pedestrians in the image; the pedestrian segmentation mask is encoded using an encoder to obtain a binary mask.

[0115] Step 3: Based on the Open Source Pose Estimation (OpenPose) algorithm, extract the human body posture key points in the image and obtain the human body posture key point map.

[0116] In this step, the OpenPose algorithm is used to detect the key points of human posture in the image and obtain the key point map of human posture. Among them, the key points of human posture refer to the location points of representative parts of the human skeleton structure, which can concisely and effectively describe the posture and movement of the human body.

[0117] S2012, based on the pre-acquired image of the clothing to be changed, acquiring a target clothing image that fits the human body in the image;

[0118] In this step, in order to replace the clothing in the image with uniform clothing, it is necessary to obtain a target clothing image that fits the human body posture in the image based on the pre-prepared clothing image to be replaced.

[0119] In a specific implementation, the target clothing image can be obtained in the following manner:

[0120] Step 1: Based on the pre-acquired image of the clothing to be changed, a Thin Plate Spline (TPS) algorithm is used to obtain a preliminary deformed image of the clothing to be changed; wherein the image of the clothing to be changed is an image of the clothing in a flat state;

[0121] In this step, when the image to be changed is an image of the clothes laid flat, the TPS algorithm can be used to adjust the original image to be changed to obtain a preliminary deformed image of the clothes to be changed. It can be expressed by the following formula:

[0122]

[0123] In the above formula, Indicates the image to be replaced; A diagram showing the initial transformation of the garment to be changed.

[0124] Step 2: Based on the masked image, the initially deformed clothing image to be changed, and the human body posture key point map, a U-shaped Network (Unet) algorithm is used to obtain a target clothing image that fits the human body in the image.

[0125] In this step, based on the masked image, the initially deformed clothing image to be changed, and the human body posture key point map, the Unet algorithm is used to make a fine prediction of the initially deformed clothing image to be changed, and obtain the target clothing image that fits the human body in the image. It can be expressed by the following formula:

[0126]

[0127] In the above formula, represents the target image; Represents the key point graph of human body posture; Represents the masked image.

[0128] S2013, encoding is performed according to the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing.

[0129] In this step, an encoder is used to encode the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing. The role of encoding is to transfer the image from the pixel space to the latent space to reduce the data processing difficulty of the image generation model.

[0130] S2014. Obtain a changed-dress image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map, and the encoded target clothing.

[0131] In this step, based on the binary mask obtained in the above step, the encoded masked image, the human body posture key point map and the encoded target clothing, an image generation model is used to obtain the image after the image is changed.

[0132] In a specific implementation, obtaining a changed-dress image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing includes:

[0133] Step 1: Get the preset text prompt information.

[0134] In this step, the text prompt information is used to prompt the image generation model to change the clothes of the pedestrians in the image.

[0135] Step 2: Input the text prompt information, random noise, binary mask, encoded masked image, human posture key point map and encoded target clothing into the inpainting pipeline of the stable diffusion model for dressing process to obtain the corresponding dressed image.

[0136] In this step, the image generation model used is the Stable Diffusion model; random noise is a fixed input to the inpainting pipeline of the Stable Diffusion model. It is worth noting that the size of the human pose key point map is the size of the image in the original clothing-changing pedestrian re-identification dataset. Before inputting it into the inpainting pipeline of the Stable Diffusion model, it is necessary to resize it to be consistent with the encoded image to obtain the adjusted human pose key point map, where the encoded image refers to the encoded masked image.

[0137] It should be understood that the dimensions of the original spatial input of the stable diffusion model inpainting pipeline are ,This scheme extends the inpainting pipeline of the stable diffusion model by connecting the input ,adjusted human pose key point map with the encoded masked image.

[0138] The input to the inpainting pipeline of the stable diffusion model can be expressed as follows:

[0139]

[0140] In the above formula, Represents the input of the stable diffusion model inpainting pipeline; represents a binary mask; The time step is The noise vector of represents the encoded masked image; represents the target garment to be encoded; Represents the adjusted human body posture key point map; Represents the dimension of the input data, specifically, is the original input dimension of the model; is the dimension of the adjusted human pose keypoint map; The dimension of the code, Indicates the height of the encoded image. Indicates the width of the encoded image.

[0141] Based on the description of steps S2011 to S2014, and the description of various implementation methods in each step, Figure 4 A schematic diagram of the principle of a specific method for generating a changed-clothing image provided in this application, such as Figure 4 As shown ( Figure 4 The original image in the figure refers to the image in the original clothing-changing pedestrian re-identification dataset). For each original image, the method first obtains the pedestrian segmentation image based on the SCHP algorithm, obtains the mask based on the segmentation mask image, and then obtains the masked image and the binary mask based on the mask; then, based on the OpePose algorithm, the human posture key point map is extracted from the original image; according to the obtained human posture key point map, the masked image and the pre-acquired clothing to be changed, the target clothing map that fits the human posture is obtained; the encoder is used to encode the target clothing map and the masked image respectively to obtain the encoded target clothing and the encoded masked image; finally, the text prompt information, random noise, the encoded target clothing, the encoded masked image and the binary mask are input into the inpainting pipeline of the stable diffusion model to obtain the image after clothing change.

[0142] In this implementation, by adopting the method of steps S2021 to S2024, the masked image is first obtained based on the pedestrian segmentation mask to ensure that the pedestrian features other than clothing in the original image (i.e., the image in the original clothing-changing pedestrian re-identification dataset) can be retained; and by obtaining the key point map of the human body posture, it is ensured that the posture of the pedestrian in the image after the clothing change is consistent with the posture of the pedestrian in the original image; and then by adjusting the image of the clothes to be changed to obtain the target clothes, the clothes of the pedestrian in the image after the clothing change are matched with the human body posture. Based on this, the method provided in the present implementation can effectively reduce the problems of texture blurring and posture distortion in the images after clothing change generated in the prior art, and can provide high-quality images after clothing change for model training.

[0143] S202: Merge the similar clothing pedestrian dataset and the original clothing-changed pedestrian re-identification dataset to obtain an expanded dataset.

[0144] In this step, it is necessary to merge the constructed similar clothing pedestrian dataset with the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset for model training.

[0145] Optionally, before merging the similar clothing pedestrian dataset and the original clothing-changed pedestrian re-identification dataset, low-quality images in the similar clothing pedestrian dataset may be removed.

[0146] In a specific implementation method, for each image after changing clothes, a feature extraction model is used to calculate the feature vectors of the image after changing clothes and its corresponding image in the original clothing-changing pedestrian re-identification dataset, and the similarity between the feature vector of the image after changing clothes and its corresponding image in the original clothing-changing pedestrian re-identification dataset is calculated; based on a similarity threshold filter, if the re-identification similarity between the image after changing clothes and its corresponding image in the original clothing-changing pedestrian re-identification dataset is lower than a preset similarity threshold, it will be removed from the similar clothing pedestrian dataset.

[0147] Among them, the feature extraction model can be a general trained Re-Id model, and this application does not limit the selection of the feature extraction model.

[0148] In addition, the preset similarity threshold can be determined based on actual conditions such as the quality requirements for the images after dressing for training, and the present application does not impose any restrictions on the size of the similarity threshold.

[0149] In this implementation, by calculating the similarity of the feature vectors of two images, low-quality images of people wearing similar clothing with problems such as human posture distortion are eliminated from the pedestrian data set, thereby successfully improving the quality of the expanded data set used for training, thereby improving the recognition accuracy of the re-identification model obtained through training.

[0150] S203: Construct a CAL model.

[0151] In this step, it is necessary to add a clothing classifier after the backbone network of the original Re-Identification (Re-Id) model to build a CAL model. The backbone network of the Re-Id model is a deep residual network (ResidualNetwork-50, ResNet50).

[0152] Among them, the CAL model is able to force the backbone network of the original Re-Id model to learn features unrelated to clothes by penalizing the original Re-Id model's prediction ability for different clothes of the same identity in the image.

[0153] S204: Train the CAL model according to the expanded data set to obtain a re-identification model.

[0154] In this step, multiple samples for model training are obtained based on the expanded data set, each sample includes a pedestrian's identity, including the pedestrian's image. Based on multiple samples and loss functions, the CAL model is iteratively trained for multiple rounds to obtain a re-identification model for re-identifying pedestrians.

[0155] It should be understood that in each iteration, the CAL model extracts features from the images in each sample and obtains the feature vector corresponding to each sample. The loss function is used to optimize the parameters of the model based on the multiple feature vectors corresponding to each identity extracted by the model during each iteration of training.

[0156] In a specific implementation, during the model training process of the CAL model, a joint loss function determined based on the Centroidloss function and the CAL loss function is selected as the loss function of the model.

[0157] Specifically, for an iterative training process, based on the identity identifiers in the samples, multiple samples corresponding to each identity identifier are determined, and then based on the multiple samples corresponding to each identity identifier and the feature vector corresponding to each sample, the intra-class centroid and inter-class centroid corresponding to each identity identifier are obtained, and the Centroid loss function corresponding to each identity identifier is determined according to the Euclidean distance between the intra-class centroid and the inter-class centroid corresponding to each identity identifier, and then the final Centroid loss function is determined based on the multiple Centroid loss functions corresponding to all identity identifiers; the Centroid loss function is linearly combined with the CAL loss function to obtain a joint loss function. Among them, the CAL loss function includes a clothing classification loss function and a multi-positive class loss function, which are used to optimize adversarial feature learning.

[0158] It should be understood that for each identity, when calculating the corresponding Centroid loss function, the overall samples are divided into two clusters, one cluster is the samples corresponding to the identity, and the other cluster is the samples not corresponding to the identity. Based on this, the intra-class centroid and inter-class centroid corresponding to each identity can be obtained.

[0159] The calculation formula of the centroid within each class corresponding to each identity can be expressed as follows:

[0160]

[0161] In the above formula, Indicates The identification of each pedestrian; Indicates that the identity is The total number of samples; Indicates The identity of the samples; Indicates The feature vector of samples; Indicates identity The corresponding centroid within the class. It should be understood that the identity The corresponding intra-class centroid is the corresponding identity marker The mean of the eigenvectors of .

[0162] The calculation formula of the centroid between classes corresponding to each identity can be expressed as follows:

[0163]

[0164] In the above formula, Indicates the total number of identity tags; Indicates identity

[0165] The corresponding centroid between classes. It should be understood that the identity The corresponding inter-class centroid is the number of all corresponding identities that are not The mean of the eigenvectors of .

[0166] Centroid loss function for each identity It can be expressed by the following formula:

[0167]

[0168] The final Centroid loss function corresponding to each identity

[0169] It can be expressed by the following formula:

[0170]

[0171] Furthermore, the joint loss function can be expressed as follows:

[0172]

[0173] In the above formula, represents the joint loss function of the CAL model; is the CAL loss function, and It is a hyperparameter used to adjust the weights of the CAL loss function and the Centroid loss function.

[0174] In this implementation, the Centroid loss function is introduced into the training of the CAL model. The Centroid loss function can minimize the intra-class distance and increase the inter-class distance, ensuring that multiple samples corresponding to the same identity identification are more compactly distributed in the feature space, effectively improving the distribution of the feature space, so that the model can focus on learning the identity characteristics of pedestrians during training, thereby enhancing the re-identification model's ability to distinguish pedestrians of different identities wearing similar clothing.

[0175] The pedestrian re-identification method based on data enhancement provided in this embodiment obtains the original clothing-changing pedestrian re-identification data set, and changes the clothing of pedestrians in multiple images in the original clothing-changing pedestrian re-identification data set into the same clothing, thereby obtaining a similar clothing pedestrian data set composed of multiple images after clothing change; then the similar clothing pedestrian data set and the original clothing-changing pedestrian re-identification data set are merged to obtain an expanded data set; then, a CAL model is constructed; finally, according to the expanded data set, the CAL model is trained to obtain a re-identification model, and based on the constructed similar clothing pedestrian data set, the generalization ability of the re-identification model in cross-clothing change scenes is successfully enhanced, thereby achieving the technical effect of improving the recognition accuracy of the model. In addition, through the method for generating images after clothing change provided by this application, and the low-similarity filtering mechanism, the quality of the images used for model training is successfully improved, and the recognition accuracy of the model is further improved. Moreover, this embodiment further enhances the ability of the re-identification model to distinguish pedestrians of different identities wearing similar clothing by introducing the Centroid loss function in the model training stage.

[0176] Figure 5 This is a schematic diagram of the structure of a pedestrian re-identification system based on data enhancement provided in Example 3 of the present application, such as Figure 5 As shown, the pedestrian re-identification system 30 based on data enhancement includes:

[0177] Electronic device 301, terminal device 302;

[0178] Among them, the electronic device 301 is used to execute the pedestrian re-identification method based on data enhancement in the above method embodiment to obtain a recognition result; the terminal device 302 is used to output the recognition result.

[0179] Specifically, the terminal device 302 may be a mobile phone, a tablet computer, a laptop computer, etc., and the present application does not limit the type of the terminal device. The user may view the recognition result on the display screen of the terminal device.

[0180] In one possible implementation, the data enhancement-based pedestrian re-identification system may also include: an image acquisition device for acquiring images in an image database in real time, which may include a wide-angle camera, a standard lens camera, a panoramic camera, a hemispherical camera, a spherical camera, a telephoto camera, and the like. This application does not impose any restrictions on the selection of the image acquisition device.

[0181] In a possible implementation, the terminal device is also used to upload the image to be identified to an electronic device.

[0182] The specific implementation process of the electronic device 301 can refer to the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0183] Furthermore, Figure 6 The system architecture of the specific pedestrian re-identification system based on data enhancement provided in the fourth embodiment of the present application is based on the above-mentioned pedestrian re-identification system based on data enhancement. The system architecture can be applied to scenes where pedestrian re-identification is required, such as large amusement parks, Figure 6 As shown, the system architecture includes: presentation layer, control layer, business logic layer, data layer and database.

[0184] Specifically, the system adopts a browser / server (B / S) architecture with separated front-end and back-end. Docker containerized deployment is used to ensure the portability and high performance of the system, so that the front-end of the system is deployed in a server local to the terminal device, and the back-end runs remotely in an Ubuntu server supported by a graphics processing unit (GPU). JSON Web Token (JWT) is used to control the permissions of the control layer, business logic layer, and data layer to ensure the security of the system and protect sensitive data in the system.

[0185] The following is a detailed introduction to the display layer, control layer, business logic layer, data layer and database:

[0186] The database is used to save and read the parameters of the re-identification model and to store the addresses of the surveillance video files stored locally. The data storage methods include MySQL database storage and local file storage.

[0187] The data layer is used to transmit data with the database based on SQLAlchemy, and prepare the pedestrian images, pedestrian videos, equipment, user accounts and other data required by the system, including camera information data, pedestrian image data, surveillance video data, and user information data. SQLAlchemy is a SQL toolkit and object-relational mapper in the Python programming language.

[0188] The business logic layer is used to read the required data based on the Flask application framework, complete the back-end logic processing of related functions, and then realize the functions of the functional modules, including user login, traffic monitoring, video trajectory tracking, stranger alarm, user information management, tourist preference analysis, image retrieval tracking and data management.

[0189] The control layer is used to obtain JSON data processed by the backend through the Application Programming Interface (API), including the execution of POST requests and GET requests;

[0190] The presentation layer is used to build an interactive user interface based on the Vue framework and feed back the results to the web browser page. It specifically includes the scripting language (JavaScript) used to control the behavior of the web page, the Hyper Text Markup Language (HTML) used to define the structure and content of the web page, and the Cascading Style Sheets (CSS) used to define the presentation style of the web page.

[0191] Furthermore, Figure 7 It is a schematic diagram of the specific functional modules in the business logic layer of the system architecture, such as Figure 7 As shown, the main functional modules in the system architecture include pedestrian flow monitoring module, intelligent search module, security module, analysis module, data management module and user management module.

[0192] Specifically, based on permission verification, the administrator can access all functional modules on the terminal device; visitors can only access some functions of the traffic monitoring module, analysis module and user management module on the terminal device.

[0193] Among them, the crowd flow monitoring module allows administrators to view park maps, queue time predictions for various amusement facilities, real-time detection of crowd flow at the entrances and exits of various facilities, and crowd flow statistics, and allows tourists to view park maps, queue time predictions for various amusement facilities, and recommend the best tour routes. It should be understood that by monitoring the crowd density of various areas in the park, combined with queue time predictions, the best tour routes are provided to tourists, and timely warnings are issued for high-density areas, effectively reducing safety hazards and improving tourists' tour experience.

[0194] The intelligent person search module can perform image-to-video pedestrian tracking and image-to-image pedestrian tracking. It should be understood that after the user uploads the image to be identified including the target pedestrian, the system can detect, track and re-identify the target pedestrian based on the data enhancement-based pedestrian re-identification method provided in this application.

[0195] It is worth noting that the YOLO (You Only Look Once) version 8 (Ultralytics YOLOv8) technology released by Ultralytics, which supports multiple hardware acceleration optimizations and is adaptable to video inputs of different resolutions, can be used as the system target detection algorithm to achieve accurate identification of target pedestrians.

[0196] The Bottom-Up Tracking by Sorting (BoT-SORT) algorithm, which can be used for pedestrian tracking tasks in large-scale and dynamic environments, can be used as the system's multi-target tracking algorithm. This algorithm improves the robustness and accuracy of tracking by comprehensively analyzing the advantages of the target's motion trajectory and appearance features. The motion compensation of the camera in dynamic scenes by this algorithm can ensure that the tracking effect is not affected by jitter. The advantage of using an improved Kalman filter state vector in this algorithm can provide higher accuracy for target prediction.

[0197] Based on this, for the intelligent search module, the system can use the above-mentioned Ultralytics YOLOv8 algorithm, BoT-SORT algorithm and the data enhancement-based pedestrian re-identification method provided in this application to retrieve surveillance video data and / or pedestrian image data, and return the trajectory video and / or image of the target pedestrian to the terminal device, as well as the tracking path drawn on the map, to achieve accurate tracking of the target pedestrian.

[0198] The security module has a stranger alarm function. It should be understood that the pedestrian re-identification method based on data enhancement provided by this application can accurately identify unauthorized persons who mistakenly enter the work area or dangerous area and trigger an alarm, significantly improving the park's security management capabilities.

[0199] The analysis module is specifically used to analyze tourists' travel preferences, and the system will make targeted activity recommendations for different tourists; the data management module has the management function of pedestrian image data, pedestrian video data, and camera data; the user management module has the functions of user information query and modification, new user information, and personal information query and modification.

[0200] The system architecture of the specific data-enhanced pedestrian re-identification system provided in this embodiment can realize accurate identification and tracking of pedestrians in indoor and outdoor scenes, across time and across clothing changes through the setting of modules such as the intelligent search module and the security module, thereby meeting the actual needs of amusement park safety management and improving the park's operating efficiency and management intelligence level.

[0201] Figure 8 This is a schematic diagram of the structure of a pedestrian re-identification device based on data enhancement provided in Example 5 of the present application, such as Figure 8 As shown, the pedestrian re-identification device 40 based on data enhancement provided in this embodiment includes:

[0202] A first acquisition unit 401 is used to acquire an image to be identified, wherein the image to be identified includes a target pedestrian;

[0203] The recognition unit 402 is used to re-recognize the target pedestrian using a re-recognition model based on the image database and the image to be recognized, and obtain a recognition result, wherein the recognition result includes at least one target image including the target pedestrian obtained from the image database;

[0204] Among them, the re-identification model is a model for re-identifying pedestrians obtained by training the CAL model with an expanded data set. The expanded data set is composed of the original pedestrian re-identification data set and the similar clothing pedestrian data set. The similar clothing pedestrian data set includes images obtained after pedestrians in multiple images in the original pedestrian re-identification data set are changed into similar clothes.

[0205] The pedestrian re-identification device 40 based on data enhancement provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.

[0206] Fig. 9 A schematic diagram of the structure of a pedestrian re-identification device based on data enhancement provided in Example 6 of the present application, such as Fig. 9 As shown, based on the above embodiment, the pedestrian re-identification device 40 based on data enhancement provided in this embodiment further includes:

[0207] The clothing changing unit 403 is used to change the clothing of pedestrians in multiple images in the original clothing changing pedestrian re-identification dataset into the same clothing, so as to obtain a similar clothing pedestrian dataset consisting of multiple clothing changing images; wherein the original clothing changing pedestrian re-identification dataset includes multiple images containing pedestrians;

[0208] An expansion unit 404 is used to merge the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain an expanded dataset;

[0209] A construction unit 405 is used to construct a CAL model;

[0210] The training unit 406 is used to train the model according to the expanded data set to obtain the re-identification model.

[0211] The filtering unit 407 is used to remove low-quality images from the similar clothing pedestrian dataset.

[0212] In a possible implementation, in the training unit 406, during the model training process of the CAL model, a joint loss function determined based on the Centroid loss function and the CAL loss function is selected as the loss function of the model.

[0213] In a possible implementation, the clothing changing unit 403 includes:

[0214] The first acquisition module is used to acquire, for each of the multiple images in the original clothing-changing pedestrian re-identification dataset, a masked image, a binary mask, and a human posture key point map corresponding to the image;

[0215] A second acquisition module is used to acquire a target clothing image that fits the human body in the image based on the pre-acquired image of the clothing to be changed;

[0216] A third acquisition module is used to encode the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing;

[0217] The fourth acquisition module is used to acquire the changed-dress image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing.

[0218] In a possible implementation manner, the first acquisition module is specifically configured to:

[0219] The image is processed based on the SCHP algorithm to generate a pedestrian segmentation mask;

[0220] According to the pedestrian segmentation mask, the masked image and the binary mask are obtained;

[0221] Based on the OpenPose algorithm, the key points of human posture in the image are extracted to obtain the key point map of human posture.

[0222] In a possible implementation manner, the second acquisition module is specifically configured to:

[0223] Based on the pre-acquired image of the clothing to be changed, the TPS algorithm is used to obtain a preliminary deformed image of the clothing to be changed; wherein the image of the clothing to be changed is an image of the clothing in a flat state;

[0224] According to the masked image, the initially deformed clothing image to be changed, and the human body posture key point map, the Unet algorithm is used to obtain the target clothing image that fits the human body in the image.

[0225] In a specific implementation, the fourth acquisition module is specifically used to:

[0226] Obtaining preset text prompt information, where the text prompt information is used to prompt pedestrians in the image to change clothes;

[0227] The text prompt information, random noise, binary mask, encoded masked image, human posture key point map and encoded target clothing are input into the inpainting pipeline of the stable diffusion model for dressing processing to obtain the dressed-up image corresponding to the image.

[0228] The pedestrian re-identification device 40 based on data enhancement provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.

[0229] Fig.10 This is a schematic diagram of the structure of an electronic device provided in Embodiment 7 of the present application. The electronic device is located in the pedestrian re-identification system based on data enhancement in Embodiment 3. Fig.10 As shown, the electronic device 50 provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the device 50 also includes a communication component 503. The processor 501, the memory 502 and the communication component 503 are connected via a bus 504.

[0230] In a specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned pedestrian re-identification method based on data enhancement.

[0231] The specific implementation process of the processor 501 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0232] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0233] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0234] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0235] The present application also provides a computer program product, including a computer program, which implements the above-mentioned pedestrian re-identification method based on data enhancement when executed by a processor.

[0236] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above-mentioned pedestrian re-identification method based on data enhancement is implemented.

[0237] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0238] An exemplary readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (Application Specific Integrated Circuits, referred to as: ASIC). Of course, the processor and the readable storage medium can also exist in the device as discrete components.

[0239] The division of units is only a logical function division, and there may be other divisions in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0240] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0241] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0242] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0243] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0244] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A pedestrian re-identification method based on data enhancement, characterized in that: include: Acquire an image to be identified, wherein the image to be identified includes a target pedestrian; Based on the image database and the image to be identified, re-identify the target pedestrian using a re-identification model to obtain a recognition result, wherein the recognition result includes at least one target image including the target pedestrian obtained from the image database; Among them, the re-identification model is a model for re-identifying pedestrians obtained by training the CAL model with an expanded data set. The expanded data set is composed of an original pedestrian re-identification data set and a pedestrian data set with similar clothing. The pedestrian data set with similar clothing includes images obtained after pedestrians in multiple images in the original pedestrian re-identification data set are changed into similar clothing.

2. The method according to claim 1, characterized in that The method further comprises: The clothing of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset is changed into the same clothing, so as to obtain the similar clothing pedestrian dataset composed of multiple images after the clothing-changing; the original clothing-changing pedestrian re-identification dataset includes multiple images containing pedestrians; Merging the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset; Constructing the CAL model; The CAL model is trained according to the expanded data set to obtain the re-identification model.

3. The method according to claim 2, characterized in that During the model training process of the CAL model, a joint loss function determined based on the Centroid loss function and the CAL loss function is selected as the loss function of the model.

4. The method according to claim 2, characterized in that: The step of changing the clothes of pedestrians in multiple images in the original clothing-changing pedestrian re-identification dataset into the same clothes comprises: For each of the multiple images in the original clothing-changing pedestrian re-identification dataset, obtain a masked image, a binary mask, and a human posture key point map corresponding to the image; Based on the pre-acquired image of the clothing to be changed, acquiring a target clothing image that fits the human body in the image; Encoding is performed according to the masked image and the target clothing image to obtain an encoded masked image and an encoded target clothing; According to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing, a changed-dress image corresponding to the image is obtained.

5. The method according to claim 4, characterized in that The step of obtaining a masked image, a binary mask and a human body posture key point map corresponding to the image includes: Processing the image based on the SCHP algorithm to generate a pedestrian segmentation mask; According to the pedestrian segmentation mask, acquiring the masked image and the binary mask; Based on the OpenPose algorithm, the human body posture key points in the image are extracted to obtain the human body posture key point map.

6. The method according to claim 4 or 5, characterized in that: The step of obtaining a target clothing image that fits the human body in the image based on the pre-acquired clothing image to be changed includes: Based on the pre-acquired image of the clothing to be changed, a TPS algorithm is used to obtain a preliminary deformed image of the clothing to be changed; wherein the image of the clothing to be changed is an image of the clothing in a flat state; According to the masked image, the initially deformed clothing image to be changed, and the human body posture key point image, the Unet algorithm is used to obtain the target clothing image that fits the human body in the image.

7. The method according to claim 4 or 5, characterized in that: The step of obtaining a dressed-up image corresponding to the image according to the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing comprises: Obtaining preset text prompt information, where the text prompt information is used to prompt a pedestrian in the image to change clothes; The text prompt information, random noise, the binary mask, the encoded masked image, the human body posture key point map and the encoded target clothing are input into the inpainting pipeline of the stable diffusion model for dressing processing to obtain the dressed-up image corresponding to the image.

8. The method according to any one of claims 2 to 5, characterized in that: Before merging the similar clothing pedestrian dataset and the original clothing-changing pedestrian re-identification dataset to obtain the expanded dataset, the method further includes: The low-quality images in the similar clothing pedestrian dataset are removed.

9. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the pedestrian re-identification method based on data enhancement as described in any one of claims 1-8.

10. A pedestrian re-identification system based on data enhancement, characterized in that: include: Electronic equipment, terminal equipment; The electronic device is used to execute the pedestrian re-identification method based on data enhancement according to any one of claims 1 to 8 to obtain a recognition result; The terminal device is used to output the recognition result.

Citation Information

Cited By

  • Image retrieval method, device and equipment

    CN120929633A