Pedestrian re-identification method and device based on diffusion model

By adopting a diffusion model-based method in pedestrian re-identification, combined with SMPL model and data enhancement technology, the identification error problem caused by changes in pedestrian geographical location and appearance characteristics in traditional methods is solved, and higher recognition accuracy and consistency are achieved.

CN119992639AInactive Publication Date: 2025-05-13AIPARK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411785793.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the pedestrian re-identification method changes in the geographical location and appearance characteristics of pedestrians, the error rate is high, the recognition gap is large, and the appearance characteristics differences caused by different camera perspectives are difficult to deal with.

Method used

The pedestrian re-identification method based on the diffusion model is adopted to identify monitoring data by constructing a diffusion model, and human pose data is constructed in combination with the SMPL model, and data from different poses and perspectives are added to the training set to enhance the recognition ability of the model.

Benefits of technology

It effectively reduces the error rate of pedestrian re-identification tasks and improves the accuracy and consistency of recognition, especially when pedestrian posture and camera perspective changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992639A_ABST
    Figure CN119992639A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian re-identification method and device based on a diffusion model. The method comprises the following steps: acquiring monitoring data and historical data of a monitoring scene; preprocessing the historical data to construct sample data; constructing a diffusion model, training the diffusion model by using the sample data, and determining a target diffusion model; and identifying the monitoring data by using the target diffusion model, and determining a pedestrian re-identification result. According to the invention, by constructing a model architecture based on a diffusion model, pedestrian re-recognition training data is subjected to data enhancement, including the posture of a pedestrian and the visual angle of a camera; data of the same pedestrian in different postures and different visual angles are added in the training set, so that the influence of different appearance features caused by different postures and visual angles on the model is reduced, and the pedestrian re-recognition capability of the model in scenes with different appearance features caused by different visual angles or postures is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent transportation technology, and in particular to a pedestrian re-identification method and device based on a diffusion model. Background Art

[0002] As the urbanization process continues to accelerate, the number of motor vehicles continues to increase, and many problems such as urban traffic congestion and parking conflicts have also arisen. At present, the main way to solve this problem is to install high-position video surveillance equipment on the side of urban roads, important traffic light intersections, etc., collect video image data of monitoring scenes, and use visual algorithms to process and analyze the data to achieve analysis of monitoring scenes, such as using pedestrian re-identification algorithms to assist in the capture of criminal suspects.

[0003] However, traditional pedestrian re-identification methods are mainly based on the fact that the tracked pedestrians move within a limited area and a certain time range, and the appearance characteristics of the pedestrians will not change significantly. However, in actual application scenarios, the tracked pedestrians usually undergo large geographical changes, and after a certain period of time, the appearance characteristics and posture of the pedestrians change significantly. In addition, due to the different installation angles of roadside surveillance cameras, the appearance characteristics of the same pedestrian captured have certain differences, which leads to problems such as high error rate and large recognition gap in pedestrian re-identification tasks. Summary of the invention

[0004] In view of this, an embodiment of the present invention provides a method and device for pedestrian re-identification based on a diffusion model, which solves the problems of high error rate and large recognition gap in pedestrian re-identification tasks in the prior art.

[0005] According to a first aspect, an embodiment of the present invention provides a method for pedestrian re-identification based on a diffusion model, comprising:

[0006] Obtain monitoring data and historical data of monitoring scenarios;

[0007] Preprocessing the historical data to construct sample data;

[0008] Constructing a diffusion model, training the diffusion model using the sample data, and determining a target diffusion model;

[0009] The target diffusion model is used to identify the monitoring data to determine a pedestrian re-identification result.

[0010] In combination with the first aspect, in the first implementation of the first aspect, the acquisition of monitoring data and historical data of the monitoring scene includes: the monitoring data and the historical data are video image data, and the video image data of each monitoring scene and the corresponding acquisition device information are acquired using existing acquisition devices in the monitoring scene.

[0011] In combination with the first implementation of the first aspect, in the second implementation of the first aspect, preprocessing the historical data to construct sample data includes:

[0012] Performing pixel conversion on the video image data corresponding to the historical data to determine conversion data;

[0013] Adding Gaussian noise to the converted data for data refinement to determine noise data;

[0014] A denoising operation is performed on the noise data to determine sample data.

[0015] In combination with the second implementation of the first aspect, in the third implementation of the first aspect, before the denoising operation is performed on the noise data, the method further includes:

[0016] Construct human posture data based on SMPL model;

[0017] Using the human body posture data and the acquisition device information, aligning the denoised data;

[0018] The pedestrian ID of the human posture data is added to the aligned data to determine the sample data.

[0019] In combination with the second implementation of the first aspect, in the fourth implementation of the first aspect, the construction of the diffusion model includes: an encoder and a denoising network,

[0020] The encoder is used to extract encoding features from the sample data;

[0021] The denoising network is a denoising network based on UNet, and the denoising network includes a downsampling module, a residual module and an upsampling module, which are used to predict the encoding features.

[0022] In combination with the fourth implementation of the first aspect, in the fifth implementation of the first aspect, coding feature extraction is performed by the following formula:

[0023] z=ε(x)

[0024] Among them, X represents the input sample data, ε represents the encoder of the diffusion model, and Z represents the encoding feature.

[0025] In combination with the first aspect, in a sixth implementation of the first aspect, the using the sample data to train the diffusion model to determine the target diffusion model includes:

[0026] Using a loss function to train the diffusion model to determine a target diffusion model;

[0027] The loss function is:

[0028]

[0029] Among them, Z t represents the encoding features in the tth stage of adding noise, and C represents the conditional information in the process of generating the image;

[0030] The conditional information is added to the image generation process by utilizing the CLIP text encoder.

[0031] The pedestrian re-identification method based on the diffusion model provided by the embodiment of the present invention takes into account the pedestrian's posture and camera perspective issues at the same time, and performs data enhancement on the pedestrian re-identification training data, including the pedestrian's posture and camera perspective, by constructing a model architecture based on the diffusion model; by adding data of the same pedestrian in different postures and different perspectives to the training set, the influence of different appearance features caused by different postures and perspectives on the model is reduced, and the model's ability to re-identify pedestrians in scenarios with different appearance features due to different perspectives or postures is enhanced.

[0032] According to a second aspect, an embodiment of the present invention provides a diffusion model-based pedestrian re-identification device, comprising:

[0033] A first processing module, used to obtain monitoring data and historical data of the monitoring scene;

[0034] A second processing module is used to pre-process the historical data to construct sample data;

[0035] A third processing module is used to construct a diffusion model, train the diffusion model using the sample data, and determine a target diffusion model;

[0036] The fourth processing module is used to use the target diffusion model to identify the monitoring data and determine the pedestrian re-identification result.

[0037] According to the third aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the diffusion model-based pedestrian re-identification method described in the first aspect or any one of the embodiments of the first aspect by executing the computer instructions.

[0038] According to a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the pedestrian re-identification method based on the diffusion model described in the first aspect or any one embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0040] Figure 1 is a flow chart of a method for pedestrian re-identification based on a diffusion model according to an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of a pedestrian re-identification device based on a diffusion model according to an embodiment of the present invention;

[0042] Figure 3 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0044] In recent years, with the continuous acceleration of urbanization, the number of motor vehicles has continued to increase, and many problems such as urban traffic congestion and parking conflicts have also arisen. To this end, relying on artificial intelligence algorithms, cloud service platforms, intelligent hardware devices, and edge computing devices, an intelligent urban traffic management system has been realized, which can collect, process, and feedback traffic information in real time and accurately.

[0045] At present, the analysis of the monitoring scene is mainly achieved by installing high-position video surveillance equipment on the side of urban roads, important traffic light intersections, etc., collecting video image data of the monitoring scene, and using visual algorithms to process and analyze the data. At present, the data collected by these monitoring devices can not only be used for traffic management, such as capturing evidence and issuing warnings for traffic violations such as running red lights and speeding, guiding roadside parking and recording parking spaces, and updating, predicting and publishing traffic congestion conditions in real time; in addition, more intelligent management of the entire city can be achieved, such as using pedestrian re-identification algorithms to assist in the capture of criminal suspects.

[0046] At present, the task of pedestrian re-identification is to identify the same person in different surveillance images. Face recognition and pedestrian appearance features are mainly used to re-identify pedestrians. Traditional pedestrian re-identification methods are mainly based on the fact that the tracked pedestrians move in a limited area and a certain time range, and the pedestrian's appearance features will not change significantly. However, in actual application scenarios, the tracked pedestrians usually have a large geographical change, and after a certain period of time, the pedestrian's appearance features and posture have changed significantly. In addition, due to the different installation angles of roadside surveillance cameras, the appearance features of the same pedestrian captured have certain differences, resulting in the pedestrian re-identification task still facing great challenges;

[0047] In addition, in supervised person re-identification tasks, the same pedestrian in different camera perspectives needs to be labeled as the same pedestrian ID. Data labeling across cameras is very difficult and requires a lot of manual work to collect, screen and label data, which consumes a lot of human and material resources. At present, in response to the problem of data collection and labeling, a method of using generative adversarial networks to perform data augmentation is proposed. However, the generative adversarial model relies on existing person re-identification data for data enhancement, and has not effectively solved the problem of pedestrian posture diversity. Although the same pedestrian has the same posture in different images, when the camera perspectives of the images are different, the pedestrian's appearance features will still be quite different, resulting in incorrect matching in pedestrian re-identification.

[0048] In this embodiment, a pedestrian re-identification method based on a diffusion model is provided, which can be used in electronic devices such as computers, mobile phones, tablet computers, etc. Figure 1 is a flowchart of a method for pedestrian re-identification based on a diffusion model according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:

[0049] S11, obtaining monitoring data and historical data of the monitoring scene.

[0050] Among them, by using cameras used in high-position video surveillance equipment or video pole surveillance equipment, pedestrians are photographed from different perspectives. The monitoring scenes include major urban roads, intersections, roadside parking areas, highways, key areas such as the entrances of schools and hospitals, etc., covering the video image data of vehicles, pedestrians and other targets in different locations and postures in the above different scenes. For pedestrian images (monitoring data and historical data), it includes pedestrian image data in different postures, different appearances, different time periods (daytime, nighttime), and different weather conditions.

[0051] S12, preprocessing the historical data to construct sample data. Specifically, the acquired historical data is operated to facilitate subsequent model training to ensure the efficiency and accuracy of the training.

[0052] S13, constructing a diffusion model, training the diffusion model using sample data, and determining a target diffusion model.

[0053] In practical applications, the diffusion model is constructed using existing technologies, which can be an existing deep learning network model, as long as it can achieve recognition. This embodiment is not limited to this. Then the license plate recognition model is trained using the target data to generate a target license plate recognition model to achieve accurate segmentation.

[0054] S14, using the target diffusion model to identify the monitoring data and determine the pedestrian re-identification result.

[0055] The pedestrian re-identification method based on the diffusion model provided in this embodiment takes into account the pedestrian's posture and camera perspective issues at the same time, and performs data enhancement on the pedestrian re-identification training data by constructing a model architecture based on the diffusion model, including the pedestrian's posture and camera perspective; by adding data of the same pedestrian in different postures and different perspectives to the training set, the influence of different appearance features caused by different postures and perspectives on the model is reduced, and the model's ability to re-identify pedestrians in scenarios with different appearance features caused by different perspectives or postures is enhanced.

[0056] In another embodiment, a method for pedestrian re-identification based on a diffusion model is provided, comprising the following steps:

[0057] S21, obtaining monitoring data and historical data of the monitoring scene. The detailed steps are the same as those in the above embodiment and will not be repeated here.

[0058] S22, preprocess the historical data and construct sample data.

[0059] In this embodiment, the above step S22 specifically further includes the following steps:

[0060] S221, performing pixel conversion on the video image data corresponding to the historical data to determine conversion data;

[0061] S222, adding Gaussian noise to the converted data for data refinement to determine noise data;

[0062] S223, denoising the noise data to determine the sample data. Specifically, it mainly includes two steps: the first is the forward diffusion process, in which the model starts to convert the pixels of the input original image by continuously adding Gaussian noise until the image completely becomes a pure noise state; the second is the reverse diffusion process, by continuously removing noise until the original image is restored.

[0063] Before denoising the noise data, the method also includes: constructing human posture data based on the SMPL model; aligning the denoised data using the human posture data and acquisition device information; adding the pedestrian ID of the human posture data to the aligned data to determine the sample data.

[0064] Specifically, a realistically rendered human body shape model SMPL is constructed to simulate complex human body shapes, including complex body joints in three-dimensional space; the human body shape model SMPL is derived from the article "SMPL: asked multi-person linear model", and a parametric three-dimensional model of the human body is constructed. The difference between this method and the traditional human body model is that SMPL can simulate the protrusions and depressions of human muscles during limb movement, so it can avoid the surface distortion of the human body during movement, and can more accurately depict the shape of human muscles in stretching and contraction states; the human body shape model SMPL has the following parameters to control the change of human body shape, where β represents 10 parameters such as height, weight, head-to-body ratio, etc., and θ represents the overall movement posture of the human body and the relative angles of 24 joints of the human body, where each joint has 3 degrees of freedom.

[0065] Step 4: Inject human posture and camera viewpoint condition information into the diffusion model;

[0066] Specifically, for data enhancement of pedestrian posture, previous work usually adopts the method of using pedestrian posture skeleton as a control condition to expand pedestrian posture. This method has certain defects. The pedestrian skeleton lacks depth information. When projected onto a two-dimensional image, there will be certain ambiguity in inferring the camera viewpoint information based on the skeleton information. In the absence of depth information, the model cannot determine whether the camera is located above the pedestrian or the person is very short. Specifically, the human body shape model SMPL is used to model the shape of the pedestrian. From this human body model, a 2D representation of the 3D human body, such as a depth map containing the camera viewpoint information, can be extracted by rendering the model.

[0067] The human posture and camera viewpoint are used as conditions for spatial alignment with the generated output image, and are processed by a posture guidance network respectively, and then they are spliced ​​in the channel dimension; for the human posture, the human skeleton is used as information, specifically represented as s∈R HxWx3 , for the camera viewpoint, the depth map is used as information, specifically expressed as d∈R HxWx1 ; For pedestrian ID, unlike the condition that it is consistent with the output in space, the pedestrian ID is not necessarily consistent with the output.

[0068] Specifically, a reference U-Net was constructed, which has the same structure as the above-mentioned denoising U-Net, and the weights of the U-Net were initialized with a pre-trained stable diffusion model; in order to inject the ID information of the image into the denoising U-Net, the image was first input into the reference U-Net; then, the self-attention mechanism was used to share the ID information with the denoising U-Net.

[0069] S23, constructing a diffusion model, training the diffusion model using sample data, and determining a target diffusion model.

[0070] In this embodiment, the above step S23 specifically further includes:

[0071] S231, constructing a diffusion model, including: an encoder and a denoising network, the encoder is used to extract encoding features of sample data; the denoising network is a denoising network based on UNet, the denoising network includes a downsampling module, a residual module and an upsampling module, which are used to predict the encoding features.

[0072] S232, training the diffusion model using the loss function to determine the target diffusion model;

[0073] The loss function is:

[0074]

[0075] Among them, Z t represents the encoded features in the tth stage where noise is added, and C represents the conditional information in the process of generating the image; the conditional information is added to the image generation process by utilizing the CLIP text encoder.

[0076] In this embodiment, in order to generate images with different camera angles and different human body postures, the present invention uses a diffusion model based on this information to achieve this. In order to avoid the diffusion model from generating unreasonable data during data generation, a stable diffusion model (Stable Diffusion) pre-trained based on a large amount of data is used to generate data;

[0077] Specifically, the diffusion model is a type of generative model. The purpose of the generative model is to learn to generate new data given some training data. The basic principle of the diffusion model is to start with random noise and gradually refine it through multiple steps until the output image appears.

[0078] Specifically, the diffusion model mainly includes two steps: the first is the forward diffusion process, in which the model begins to transform the pixels of the input original image by continuously adding Gaussian noise until the image completely becomes a pure noise state; the second is the reverse diffusion process, which continuously removes noise until the original image is restored;

[0079] Specifically, the Stable Diffusion model (StableDiffusion) based on a large amount of data pre-training adopts a more stable, controllable and efficient method to generate high-quality images. In terms of training data, the Stable Diffusion model is trained on the Laion-2B-en dataset, which contains 2.32 billion images and English control texts. It uses more training data and also uses data screening to improve the quality of sample data. In the text encoding process, the Stable Diffusion model performs the denoising process in the latent space of the self-encoder, which reduces the computational cost compared with denoising in the pixel space. It is specifically expressed as:

[0080] z=ε(x)

[0081] Among them, X represents the input sample data, ε represents the encoder of the diffusion model, and Z represents the encoding feature.

[0082] During the training process, the stable diffusion model uses a denoising network based on UNet for model learning. The denoising UNet model consists of three parts: a downsampling module, a residual module, and an upsampling module. Each module consists of multiple convolutional layers, self-attention layers, and cross-attention layers. Specifically, the denoising UNet model predicts the potential feature representation after the above encoding, and the output is represented as ∈. The specific training loss function is expressed as:

[0083]

[0084] Among them, Z t represents the encoded features in the tth stage with added noise, C represents the conditional information in the process of generating the image, specifically, CLIP, Contrastive Language-Image Pre-training, is a multimodal (text and image) pre-training model that embeds text and images into a common semantic space.

[0085] S24, using the target diffusion model to identify the monitoring data and determine the pedestrian re-identification result.

[0086] A pedestrian re-identification method based on a diffusion model is provided in this embodiment, which takes into account the pedestrian's posture and camera perspective. By constructing a model architecture based on the diffusion model, data enhancement is performed on the pedestrian re-identification training data, including the pedestrian's posture and camera perspective. By adding data of the same pedestrian in different postures and different perspectives to the training set, the influence of different appearance features caused by different postures and perspectives on the model is reduced, and the model's ability to re-identify pedestrians in scenarios with different appearance features due to different perspectives or postures is enhanced.

[0087] In this embodiment, a pedestrian re-identification device based on a diffusion model is also provided. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0088] The present invention discloses a pedestrian re-identification device based on a diffusion model, such as Figure 2 As shown, including:

[0089] The first processing module 01 is used to obtain monitoring data and historical data of the monitoring scene; the detailed operation process is detailed in the corresponding steps in the above method embodiment;

[0090] The second processing module 02 is used to pre-process the historical data and construct sample data; the detailed operation process is detailed in the corresponding steps in the above method embodiment;

[0091] The third processing module 03 is used to construct a diffusion model, train the diffusion model using sample data, and determine a target diffusion model; the detailed operation process is detailed in the corresponding steps in the above method embodiment;

[0092] The fourth processing module 04 is used to identify the monitoring data using the target diffusion model and determine the pedestrian re-identification result; the detailed operation process is detailed in the corresponding steps in the above method embodiment.

[0093] The present invention also provides an electronic device. Figure 3 , Figure 3 is a schematic diagram of the structure of an electronic device provided by an optional embodiment of the present invention, such as Figure 3 As shown, the electronic device may include: at least one processor 601, such as a CPU (Central Processing Unit), at least one communication interface 603, a memory 604, and at least one communication bus 602. Among them, the communication bus 602 is used to realize the connection and communication between these components. Among them, the communication interface 603 may include a display screen (Display), a keyboard (Keyboard), and the optional communication interface 603 may also include a standard wired interface and a wireless interface. The memory 604 may be a high-speed RAM memory (Random Access Memory, volatile random access memory) or a non-volatile memory (non-volatile memory), such as at least one disk storage. The memory 604 may optionally be at least one storage device located away from the aforementioned processor 601. The memory 604 stores the application program, and the processor 601 calls the program code stored in the memory 604 to execute any of the above method steps.

[0094] The communication bus 602 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The communication bus 602 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0095] Among them, the memory 604 may include a volatile memory (English: volatile memory), such as a random access memory (English: random-access memory, abbreviated: RAM); the memory may also include a non-volatile memory (English: non-volatile memory), such as a flash memory (English: flash memory), a hard disk drive (English: hard disk drive, abbreviated: HDD) or a solid-state drive (English: solid-state drive, abbreviated: SSD); the memory 604 may also include a combination of the above types of memory.

[0096] The processor 601 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and a NP.

[0097] The processor 601 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0098] Optionally, the memory 604 is also used to store program instructions. The processor 601 can call the program instructions to implement the method shown in the embodiment of the present application.

[0099] The embodiment of the present invention further provides a non-transitory computer storage medium, which stores computer executable instructions, and the computer executable instructions can execute the method in any of the above method embodiments. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory (Flash Memory), a hard disk (HDD) or a solid-state drive (SSD), etc.; the storage medium can also include a combination of the above types of memory.

[0100] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations are all within the scope defined by the appended claims.

Claims

1. A person re-identification method based on a diffusion model, characterized in that: include: Obtain monitoring data and historical data of monitoring scenarios; Preprocessing the historical data to construct sample data; Constructing a diffusion model, training the diffusion model using the sample data, and determining a target diffusion model; The target diffusion model is used to identify the monitoring data to determine a pedestrian re-identification result.

2. The method according to claim 1, characterized in that The acquiring of monitoring data and historical data of the monitoring scene includes: the monitoring data and the historical data are video image data, and the existing acquisition equipment in the monitoring scene is used to acquire the video image data of each monitoring scene and the corresponding acquisition equipment information.

3. The method according to claim 2, characterized in that The preprocessing of the historical data to construct sample data includes: Performing pixel conversion on the video image data corresponding to the historical data to determine conversion data; Adding Gaussian noise to the converted data for data refinement to determine noise data; A denoising operation is performed on the noise data to determine sample data.

4. The method according to claim 3, characterized in that Before performing the denoising operation on the noise data, the method further includes: Construct human posture data based on SMPL model; Using the human body posture data and the acquisition device information, aligning the denoised data; The pedestrian ID of the human posture data is added to the aligned data to determine the sample data.

5. The method according to claim 3, characterized in that: The diffusion model is constructed, including: an encoder and a denoising network, The encoder is used to extract encoding features from the sample data; The denoising network is a denoising network based on UNet, and the denoising network includes a downsampling module, a residual module and an upsampling module, which are used to predict the encoding features.

6. The method according to claim 5, characterized in that The encoding feature is extracted by the following formula: z=ε(x) Among them, X represents the input sample data, ε represents the encoder of the diffusion model, and Z represents the encoding feature.

7. The method according to claim 1, characterized in that The step of training the diffusion model using the sample data to determine a target diffusion model includes: Using a loss function to train the diffusion model to determine a target diffusion model; The loss function is: Among them, Z t represents the encoding features in the tth stage of adding noise, and C represents the conditional information in the process of generating the image; The conditional information is added to the image generation process by utilizing the CLIP text encoder.

8. A pedestrian re-identification device based on a diffusion model, characterized in that: include: A first processing module, used to obtain monitoring data and historical data of the monitoring scene; A second processing module is used to pre-process the historical data to construct sample data; A third processing module is used to construct a diffusion model, train the diffusion model using the sample data, and determine a target diffusion model; The fourth processing module is used to use the target diffusion model to identify the monitoring data and determine the pedestrian re-identification result.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on diffusion model

    CN116246307A

  • Virtual anchor whole-body video generation method and system based on diffusion model

    CN117979115A