A pedestrian re-identification method and related device

By extracting the local, global and semantic features of ground monitoring images and drone images, and combining multi-scale and multi-view training data, the problem of pedestrian re-identification in drone and ground monitoring scenarios is solved, and more efficient and robust pedestrian re-identification performance is achieved.

CN119832601BActive Publication Date: 2025-06-20CGN WIND POWER CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510302344.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing pedestrian re-identification technology is difficult to directly apply in drone and ground monitoring scenarios, especially when the perspective angle, scale and attitude changes greatly, and a single feature descriptor is difficult to achieve ideal re-identification performance.

Method used

By obtaining the ground monitoring image of the target pedestrian and the image set to be identified, the pedestrian re-identification model is used to extract local features, global features and semantic features, and combining multi-scale and multi-view training data to determine the drone image of the target pedestrian.

Benefits of technology

It improves the performance and robustness of pedestrian re-identification, can provide better re-identification performance when facing problems such as occlusion, and provides more context information through drone images, enhancing the model's ability to discriminate pedestrians.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832601B_ABST
    Figure CN119832601B_ABST
Patent Text Reader

Abstract

The present application discloses a pedestrian re-identification method and related device, which relates to the technical field of image processing, and includes: obtaining a ground surveillance image of a target pedestrian and a set of drone images to be identified, and respectively extracting local features, global features and semantic features of the ground surveillance image and each drone image by using a pedestrian re-identification model, so as to determine the drone image of the target pedestrian from the set of drone images to be identified. The present application trains a pedestrian re-identification model based on multi-scale and multi-view training data. The multi-scale enables the model to capture more global and local detailed features and environmental context features of pedestrians, enhancing the discriminative ability of the model for pedestrians. The multi-view enables the model to provide better re-identification performance when facing problems such as occlusion. At the same time, the drone images can collect broader environmental information, thereby being able to provide more context information for the pedestrian re-identification process, and further improving the re-identification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a pedestrian re-identification method and related devices. Background Art

[0002] Pedestrian re-identification (Person Re-identification, ReID) technology is an important technology in the field of computer vision. Its core lies in using computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. This technology is widely used in fields such as video surveillance, security, and intelligent transportation, and is of great significance for improving public safety and other aspects.

[0003] The current mainstream pedestrian re-identification methods mainly include: detection methods based on traditional image pedestrian features and end-to-end feature extraction and matching methods based on deep learning.

[0004] The detection methods based on traditional image pedestrian features mainly use information such as the outer ring features, shape features, and gait features of pedestrians for modeling and matching. A typical approach is to extract the feature descriptors of pedestrian images, and then measure the similarity between different pedestrian descriptors through distance metrics or learned similarity functions. This method has relatively good pedestrian re-identification effects for images with limited perspective changes. However, in the UAV scenario, due to large changes in perspective, scale, and pose, it is difficult for a single feature descriptor to obtain ideal re-identification performance.

[0005] The end-to-end feature extraction and matching methods based on deep learning train a pedestrian re-identification model through large-scale training data, and automatically discover high-level pedestrian semantic features based on this pedestrian re-identification model, thereby improving the performance and robustness of re-identification. However, the training data of the pedestrian re-identification model is often captured by surveillance cameras, with a low shooting height and a stationary lens, while UAVs often capture pedestrians during high-altitude movement. The data distributions in the two scenarios are quite different, resulting in the difficulty of directly applying the pedestrian re-identification model trained in the classical scenario to the UAV scenario. Summary of the Invention

[0006] In view of the above problems, this application provides a pedestrian re-identification method and related devices to achieve the purpose of combining UAVs with ground surveillance and accurately performing pedestrian re-identification. The specific solutions are as follows:

[0007] The first aspect of this application provides a pedestrian re-identification method, including:

[0008] Obtain the ground surveillance image of the target pedestrian and the set of UAV images to be recognized, where the set of UAV images to be recognized contains multiple UAV images;

[0009] Use the person re-identification model to extract the local features, global features, and semantic features of the ground surveillance image and each of the drone images respectively, so as to determine the drone image of the target pedestrian from the set of drone images to be identified based on the extracted features;

[0010] Among them, the person re-identification model uses the training ground surveillance images and the training drone image set with labeled training pedestrian labels as training data. Each drone image included in the training drone image set has a perspective difference and a scale difference from the training ground surveillance image. The global features represent the overall appearance information of the pedestrian and the information of the pedestrian and the surrounding environment. The local features represent the detailed information of different regions of the pedestrian's body. The semantic features represent the information of the invariant attributes on the pedestrian.

[0011] In a possible implementation, for any one of the ground surveillance image and each of the drone images, the process of extracting the local features of the image includes:

[0012] Segment the image by body parts to obtain a plurality of segmented sub-images;

[0013] Use a feature extraction network to extract the image features of the plurality of sub-images respectively;

[0014] Perform attention-based feature enhancement processing on the image features of the plurality of sub-images respectively to obtain the enhanced features corresponding to the plurality of sub-images respectively;

[0015] Use the enhanced features corresponding to the plurality of sub-images respectively as the local features of the image.

[0016] In a possible implementation, for any one of the ground surveillance image and each of the drone images, the process of extracting the global features of the image includes:

[0017] Use a plurality of cascaded encoding layers to perform global feature extraction processing on the image to obtain the global features of the image;

[0018] Among them, the process of using any one of the encoding layers to perform global feature extraction processing includes:

[0019] Obtain the first feature map input to the encoding layer. Among them, if the encoding layer is the first encoding layer, the first feature map is the initial feature map obtained by performing preliminary encoding on the sequence of image blocks included in the image. If the encoding layer is not the first encoding layer, the first feature map is the second feature map output by the previous encoding layer;

[0020] Process the first feature map based on the multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to each of the multiple attention heads;

[0021] Perform inverse wavelet transform on the downsampled feature map to obtain a first reconstructed feature map and a second reconstructed feature map, where the downsampled feature map is the feature map obtained during the process of processing the first feature map based on the multi-head attention mechanism integrated with wavelet transform;

[0022] Obtain an attention summary map based on the query vectors and key-value pairs corresponding to each of the multiple attention heads, the first reconstructed feature map, and the second reconstructed feature map;

[0023] Obtain a transition feature map based on the attention summary map and the first feature map;

[0024] Obtain the second feature map output by the encoding layer based on a multi-layer perceptron and the transition feature map, where the second feature map output by the last encoding layer is used to obtain the global feature of the image.

[0025] In a possible implementation, the process of processing the first feature map based on the multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to each of the multiple attention heads includes:

[0026] Normalize the first feature map to obtain a normalized third feature map;

[0027] Perform a linear transformation on the third feature map using an embedding matrix to obtain a fourth feature map;

[0028] Decompose the fourth feature map into four wavelet subbands through discrete wavelet transform;

[0029] Concatenate the four wavelet subbands along the channel dimension, and then perform convolution processing on the concatenation result to obtain the downsampled feature map;

[0030] Perform multiple groups of linear transformations on the downsampled feature map to obtain the query vectors and key-value pairs corresponding to each of the multiple attention heads.

[0031] In a possible implementation, the process of preliminary encoding includes:

[0032] Encode each image block in the image block sequence included in the image into a vector of a preset length to obtain the vector corresponding to each image block;

[0033] Perform position encoding on each image block based on the position of each image block in the image to obtain the position encoding result corresponding to each image block;

[0034] Based on the vectors and position encoding results corresponding to all the image patches included in the image patch sequence, the initial feature map is obtained.

[0035] In a possible implementation, for any one of the ground surveillance images and each of the drone images, the process of extracting the semantic features of the image includes:

[0036] Performing convolution processing and activation processing on the global features of the image to obtain a fifth feature map corresponding to the image, where the convolution dimension of a convolution function used in the convolution processing includes the number of invariant attributes, and the activation processing is used to enhance the attention to the regions where the invariant attributes are located in the image;

[0037] Processing the global features of the image using a preset attention mechanism to obtain an attribute attention map corresponding to the image;

[0038] Based on the attribute attention map and the fifth feature map corresponding to the image, the semantic features of the image are obtained.

[0039] In a possible implementation, the loss function of the person re-identification model is a joint loss function composed of a triplet loss function, a metric distillation loss function, and a cross-entropy loss function.

[0040] The second aspect of this application provides a person re-identification device, including:

[0041] An image acquisition module, configured to acquire a ground surveillance image of a target person and a set of drone images to be recognized, where the set of drone images to be recognized includes multiple drone images;

[0042] A person re-identification module, configured to use a person re-identification model to extract the local features, global features, and semantic features of the ground surveillance image and each of the drone images respectively, so as to determine the drone image of the target person from the set of drone images to be recognized based on the extracted features;

[0043] Wherein, the person re-identification model uses the training ground surveillance images and the set of training drone images with labeled training person labels as training data, each drone image included in the set of training drone images has a perspective difference and a scale difference from the training ground surveillance image, the global features represent the overall appearance information of the person and the information of the person and the surrounding environment, the local features represent the detailed information of different regions of the person's body, and the semantic features represent the information of the invariant attributes on the person.

[0044] In a third aspect of the present application, a computer program product is provided, including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement the pedestrian re-identification method according to the first aspect or any implementation manner of the first aspect as described above.

[0045] In a fourth aspect of the present application, an electronic device is provided, including at least one processor and a memory connected to the processor, wherein:

[0046] The memory is used to store a computer program;

[0047] The processor is used to execute the computer program so that the electronic device can implement the pedestrian re-identification method according to the first aspect or any implementation manner of the first aspect as described above.

[0048] In a fifth aspect of the present application, a computer storage medium is provided. The storage medium carries one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement the pedestrian re-identification method according to the first aspect or any implementation manner of the first aspect as described above.

[0049] By means of the above technical solutions, for the pedestrian re-identification method provided in the present application, a ground monitoring image of a target pedestrian and a set of drone images to be identified are acquired, and local features, global features, and semantic features of the ground monitoring image and each drone image are respectively extracted by using a pedestrian re-identification model, so as to determine the drone image of the target pedestrian from the set of drone images to be identified based on the extracted features. The present application trains the pedestrian re-identification model based on multi-scale and multi-view training data. The multi-scale training data enables the pedestrian re-identification model to capture more global and local detailed features and environmental context features of pedestrians, enhancing the discriminative ability of the model for pedestrians. The multi-view training data enables the pedestrian re-identification model to provide better re-identification performance when facing problems such as occlusion. At the same time, drone images can collect broader environmental information, so as to provide more context information for the pedestrian re-identification process, and thus improve the re-identification performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In combination with the drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original elements and elements are not necessarily drawn to scale.

[0051] Figure 1 It is a schematic diagram of a system architecture provided by the present application;

[0052] Figure 2 It is an optional hardware structure schematic diagram of the terminal 100 provided by the present application;

[0053] Figure 3 Schematic diagram of the structure of a server 200 provided for this application;

[0054] Figure 4 Schematic flow diagram of a person re-identification method provided for this application;

[0055] Figure 5 Schematic diagram of the structure of a local feature extraction module for extracting local features of an image;

[0056] Figure 6 Schematic diagram of the structure of a ResNet50 feature extraction network;

[0057] Figure 7 Schematic diagram of the structure of the first bottleneck block in ResNet50;

[0058] Figure 8 Schematic diagram of the structure of the second bottleneck block in ResNet50;

[0059] Figure 9 Schematic diagram of the structure of an encoding layer;

[0060] Figure 10 Schematic diagram of the structure of a multi-head self-attention module;

[0061] Figure 11 Schematic diagram of the structure of an attribute parsing module;

[0062] Figure 12 Schematic diagram of the structure of a person re-identification device provided for this application;

[0063] Figure 13 Schematic diagram of the structure of an electronic device provided for this application. Detailed implementation manners

[0064] The embodiments of this application will be described below with reference to the accompanying drawings in the embodiments of this application. The terms used in the embodiments part of this application are only used to explain the specific embodiments of this application, rather than to limit this application.

[0065] The embodiments of this application will be described below with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are equally applicable to similar technical problems.

[0066] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing the embodiments of this application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0067] See Figure 1 , Figure 1 shows a schematic diagram of a system architecture. The system may include a terminal 100 and a server 200. Among them, the server 200 may include one or more servers ( Figure 1 illustrated by including one server as an example), and the server 200 may provide the method provided by the embodiments of this application for one or more terminals.

[0068] Among them, an application may be installed on the terminal 100, and the above application and web page may provide an interface. The terminal 100 may receive relevant parameters input by the user on the interface and send the above parameters to the server 200. The server 200 may obtain a processing result based on the received parameters and return the processing result to the terminal 100.

[0069] It should be understood that in some alternative implementations, the terminal 100 may also complete the action of obtaining the processing result based on the received parameters by itself without the cooperation of the server, which is not limited in the embodiments of this application.

[0070] Next, describe Figure 1 the product form of the terminal 100 in

[0071] The terminal 100 in the embodiments of this application may be a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc., and the embodiments of this application do not make any restrictions on this.

[0072] Figure 2 shows an alternative schematic diagram of the hardware structure of the terminal 100.

[0073] Reference Figure 2 As shown, the terminal 100 may include components such as a radio frequency unit 110, a memory 120, an input unit 130, a display unit 140, a camera 150 (optional), an audio circuit 160 (optional), a speaker 161 (optional), a microphone 162 (optional), a headphone jack 163 (optional), a processor 170, an external interface 180, a power supply 190, etc. Those skilled in the art can understand that Figure 2 This is merely an example of a terminal or a multifunctional device and does not constitute a limitation on the terminal or the multifunctional device. It may include more or fewer components than shown in the figure, or combine certain components, or have different components.

[0074] The input unit 130 can be used to receive input digital or character information and generate key signal inputs related to the user settings and function controls of the portable multifunctional device. Specifically, the input unit 130 may include a touch screen 131 (optional) and / or other input devices 132. The touch screen 131 can collect touch operations of the user thereon or nearby (such as operations of the user using a finger, a joint, a stylus, or any suitable object on or near the touch screen), and drive the corresponding connection device according to a pre-set program. The touch screen can detect the touch actions of the user on the touch screen, convert the touch actions into touch signals and send them to the processor 170, and can receive and execute the commands sent by the processor 170; the touch signals at least include contact coordinate information. The touch screen 131 can provide an input interface and an output interface between the terminal 100 and the user. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch screen. In addition to the touch screen 131, the input unit 130 may further include other input devices. Specifically, the other input devices 132 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0075] Among them, the input device 132 can receive input data, etc.

[0076] The display unit 140 can be used to display information input by the user or information provided to the user, various menus of the terminal 100, an interactive interface, file display, and / or the playback of any multimedia file.

[0077] The memory 120 can be used to store instructions and data. The memory 120 mainly includes an instruction storage area and a data storage area. The data storage area can store various data, such as multimedia files, texts, etc.; the instruction storage area can store software units such as an operating system, applications, and instructions required for at least one function, or subsets or extended sets thereof. It can also include a non-volatile random access memory; it provides the processor 170 with functions including managing the hardware, software, and data resources in the computing processing device, supporting control software and applications. It is also used for storing multimedia files, as well as storing running programs and applications.

[0078] The processor 170 is the control center of the terminal 100, connecting various parts of the entire terminal 100 through various interfaces and lines. By running or executing the instructions stored in the memory 120 and calling the data stored in the memory 120, it executes various functions of the terminal 100 and processes data, thereby controlling the terminal device as a whole. Optionally, the processor 170 may include one or more processing units; preferably, the processor 170 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 170 either. In some embodiments, the processor and the memory can be implemented on a single chip, and in some embodiments, they can also be separately implemented on independent chips. The processor 170 can also be used to generate corresponding operation control signals, send them to corresponding components of the computing processing device, read and process the data in the software, especially read and process the data and programs in the memory 120, so that each functional module therein executes corresponding functions, thereby controlling the corresponding components to act according to the requirements of the instructions.

[0079] Among them, the memory 120 can be used to store software codes related to the pedestrian re-identification method, and the processor 170 can execute the steps of the pedestrian re-identification method, or can also schedule other units (such as the above-mentioned input unit 130 and display unit 140) to implement corresponding functions.

[0080] The radio frequency unit 110 (optional) can be used for receiving and transmitting information or signals during a call. For example, after receiving the downlink information from the base station, it is sent to the processor 170 for processing; in addition, the uplink data designed is sent to the base station. Generally, the RF circuit includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the radio frequency unit 110 can also communicate with network devices and other devices through wireless communication. This wireless communication can use any communication standard or protocol, including but not limited to Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0081] Among them, in the embodiment of the present application, the radio frequency unit 110 can send data to the server 200 and receive the processing result sent by the server 200.

[0082] It should be understood that the radio frequency unit 110 is optional and can be replaced by other communication interfaces, such as a network interface.

[0083] The terminal 100 also includes a power supply 190 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 170 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0084] The terminal 100 also includes an external interface 180, which can be a standard Micro USB interface or a multi-pin connector, and can be used to connect the terminal 100 to other devices for communication, or to connect a charger to charge the terminal 100.

[0085] Although not shown, the terminal 100 may also include a flash, a Wireless Fidelity (WiFi) module, a Bluetooth module, sensors with different functions, etc., which will not be elaborated here. Some or all of the methods described below can be applied to the terminal 100 as Figure 2 shown.

[0086] The following describes Figure 1 the product form of the server 200 in

[0087] Figure 3 A schematic structural diagram of a server 200 is provided, as Figure 3 shown. The server 200 includes a bus 201, a processor 202, a communication interface 203, and a memory 204. The processor 202, the memory 204, and the communication interface 203 communicate with each other through the bus 201.

[0088] The bus 201 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 only a thick line is used in

[0089] to represent it, but it does not mean that there is only one bus or one type of bus.

[0090] The processor 202 can be any one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), a Micro Processor (MP), or a Digital Signal Processor (DSP).

[0091] The memory 204 can include a volatile memory, such as a Random Access Memory (RAM). The memory 204 can also include a non-volatile memory, such as a Read-Only Memory (ROM), a flash memory, a Hard Disk Drive (HDD), or a Solid State Drive (SSD).

[0092] It should be understood that the above terminal 100 and server 200 can be centralized or distributed devices, and the processors in the above terminal 100 and server 200 (such as processor 170 and processor 202) can be hardware circuits (such as Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), general-purpose processor, DSP, microprocessor or microcontroller, etc.), or a combination of these hardware circuits. For example, the processor can be a hardware system with the function of executing instructions, such as CPU, DSP, etc., or a hardware system without the function of executing instructions, such as ASIC, FPGA, etc., or a combination of the above hardware system without the function of executing instructions and the hardware system with the function of executing instructions.

[0093] This application provides a person re-identification method. The person re-identification method according to the embodiments of this application will be introduced in detail below with reference to the accompanying drawings.

[0094] Refer to Figure 4 , Figure 4 which is a schematic flowchart of a person re-identification method provided by an embodiment of this application. The method may include:

[0095] Step S401, obtain a ground surveillance image of a target pedestrian and a set of drone images to be recognized, where the set of drone images to be recognized includes multiple drone images.

[0096] In this embodiment, a ground surveillance device can be used to monitor a required area. When the target pedestrian enters the monitored area, the ground surveillance device can capture a ground surveillance image of the target pedestrian.

[0097] At the same time, this application can also capture an aerial view image by a drone to obtain a set of drone images to be recognized.

[0098] Optionally, the ground surveillance image and the multiple drone images in the set of drone images to be recognized can be synchronously acquired images or asynchronously acquired images; the ground surveillance image and the multiple drone images in the set of drone images to be recognized can have viewing angle and / or scale differences, or can have no viewing angle and scale differences; the monitored areas of the ground surveillance device and the drone can be the same or different; this application does not make specific limitations.

[0099] Optionally, the above ground surveillance image is an image that only includes the target pedestrian; optionally, the target pedestrian can include one pedestrian or multiple pedestrians. If the target pedestrian includes multiple pedestrians, the ground surveillance image in this step includes the respective ground surveillance images of multiple pedestrians.

[0100] Optionally, the above UAV image may only contain one pedestrian. It can be understood that the UAV may capture multiple pedestrians in one frame of image. For the convenience of subsequent processing, the images collected by the UAV can be first segmented according to pedestrians to obtain the above UAV image that only contains one pedestrian. Optionally, the segmentation process can be implemented by the YOLO (You Only Look Once) model.

[0101] Step S402: Use the person re-identification model to extract the local features, global features, and semantic features of each ground surveillance image and each UAV image respectively, so as to determine the UAV image of the target pedestrian from the set of UAV images to be recognized based on the extracted features.

[0102] In this embodiment, a person re-identification model can be pre-trained. Specifically, the UAV can be used to capture an overhead image to obtain a panoramic scene, and detail images can be captured synchronously or asynchronously by multiple ground surveillance devices (such as ground surveillance cameras) from different perspectives. All images are calibrated for spatio-temporal positions and the positions and identities of pedestrians are manually annotated. Finally, the paired ground-air multi-perspective image pairs are organized into training, validation, and test data sets in a specific format. Among them, the above spatio-temporal position calibration is to calibrate the image capture time and position for convenient data archiving.

[0103] Then, the training data of the person re-identification model provided in this embodiment is the training ground surveillance images and the training UAV image set labeled with training pedestrian labels. That is, in the training process, the training ground surveillance images of the training pedestrians and the training UAV image set can be input into the pre-constructed neural network model to obtain the model prediction output. Then, the model prediction output is compared with the training UAV image set labeled with training pedestrian labels (that is, through this label, it can be determined whether the pedestrian in the UAV image is a training pedestrian) and the loss is calculated. Based on the loss, the parameters of the neural network model are trained to obtain the person re-identification model.

[0104] To enable the person re-identification model to have higher recognition accuracy, optionally, each UAV image included in the training UAV image set may have a perspective difference and a scale difference from the training ground surveillance image. Among them, the perspective difference means that there is a certain difference between the UAV overhead angle and the ground surveillance angle to reflect the different characteristic features of the pedestrian target under different perspectives; the scale difference means that the scale of the pedestrian target in the UAV image is significantly smaller than that in the ground surveillance image to simulate the large-scale change in the actual situation.

[0105] Optionally, the person re-identification model is trained using the Adam (Adaptive Moment Estimation) optimizer, and the learning rate is 1e-4.

[0106] In this embodiment, the perspective difference and scale difference between image pairs can improve the robustness of the person re-identification model.

[0107] In this embodiment, the local features, global features, and semantic features of the ground surveillance image and each drone image can be extracted respectively by using the person re-identification model, that is, the local features, global features, and semantic features of the ground surveillance image are extracted, and at the same time, the local features, global features, and semantic features of each drone image also need to be extracted.

[0108] Here, the global features represent the overall appearance information of the pedestrian and the information of the pedestrian and the surrounding environment; the local features represent the detailed information of different regions of the pedestrian's body, for example, detailed features such as head features, clothing textures, and shoe styles; the semantic features represent the information of the invariant attributes on the pedestrian, such as whether the pedestrian wears a hat, whether the pedestrian carries a backpack, etc. Due to the difference in height between the drone perspective and the ground surveillance perspective, even for different images of the same pedestrian, the body posture will change greatly, but attributes such as whether wearing a hat will not change significantly. Therefore, in the process of person re-identification, the semantic features of the pedestrian image, that is, the attribute features, play an important role.

[0109] Furthermore, in this embodiment, based on the local features, global features, and semantic features of the ground surveillance image and each drone image respectively, the drone image of the target pedestrian can be determined from the set of drone images to be recognized.

[0110] For example, optionally, the local features, global features, and semantic features of the ground surveillance image can be fused to obtain the fused features of the target pedestrian; the local features, global features, and semantic features of each drone image can be fused to obtain the fused features of the pedestrians to be recognized included in each drone image. Then, the fused features of the target pedestrian are compared with the fused features of the pedestrians to be recognized included in each drone image respectively. The comparison process can be realized by calculating distances, calculating similarities, etc., so as to determine whether the pedestrians to be recognized included in each drone image are the target pedestrian through the comparison process, and then obtain the drone image of the target pedestrian.

[0111] The person re-identification method provided by this application obtains the ground surveillance image of the target person and the set of drone images to be identified, and uses the person re-identification model to extract the local features, global features, and semantic features of the ground surveillance image and each drone image respectively, so as to determine the drone image of the target person from the set of drone images to be identified based on the extracted features. This application trains the person re-identification model based on multi-scale and multi-view training data. The multi-scale training data enables the person re-identification model to capture more global and local detailed features and environmental context features of pedestrians, enhancing the discriminative ability of the model for pedestrians. The multi-view training data enables the person re-identification model to provide better re-identification performance when facing problems such as occlusion. At the same time, the drone images can collect more extensive environmental information, so as to provide more context information for the person re-identification process, thereby improving the re-identification performance.

[0112] In some embodiments of this application, the process of step S402, "using the person re-identification model to extract the local features, global features, and semantic features of the ground surveillance image and each drone image respectively", is introduced.

[0113] In this embodiment, for the ground surveillance image and each drone image, the processes of extracting local features, global features, and semantic features are the same respectively. For the convenience of introduction, the following takes any one of the ground surveillance image and each drone image as an example to introduce the process of extracting the local features, global features, and semantic features of this image.

[0114] First, the process of extracting the local features of this image includes: segmenting this image by body part to obtain multiple segmented sub-images, using a feature extraction network to extract the image features of each of the multiple sub-images, and respectively performing attention-based feature enhancement processing on the image features of each of the multiple sub-images to obtain the enhanced features corresponding to each of the multiple sub-images, and taking the enhanced features corresponding to each of the multiple sub-images as the local features of this image.

[0115] Optionally, this image can be segmented into a head sub-image, an upper body sub-image, and a lower body sub-image.

[0116] See Figure 5 , which shows a schematic structural diagram of a local feature extraction module for extracting the local features of an image. As Figure 5 , the local feature extraction module is composed of an image segmentation module, a feature extraction module, and an attention module.

[0117] Among them, the image segmentation module segments the image by body part to obtain multiple segmented sub-images. The feature extraction module uses a feature extraction network to extract the image features of each of the multiple sub-images. The attention module performs attention-based feature enhancement processing on the image features of each of the multiple sub-images respectively to obtain enhanced features corresponding to the multiple sub-images respectively, and uses the enhanced features corresponding to the multiple sub-images respectively as the local features of the image.

[0118] Optionally, the image segmentation module is implemented by a positioning layer, and the positioning layer divides different regions by performing spatial operations such as scaling, moving, and cropping. The principle is as follows formula (1):

[0119] Formula (1);

[0120] Among them, and are respectively and scaling parameters in the directions of and are respectively and translation parameters in the directions of is the original coordinate of the pixel point in the image, is the coordinate of the pixel point in the image after transformation.

[0121] Then, the position of each pixel in the segmented sub-image is determined by its original coordinate and the coordinate after transformation.

[0122] Optionally, the feature extraction module can adopt a ResNet50 feature extraction network.

[0123] See Figure 6 , which shows the structural schematic diagram of the ResNet50 feature extraction network. See Figure 7 and Figure 8 , which are respectively the structural schematic diagrams of the first bottleneck block and the second bottleneck block in ResNet50.

[0124] As Figure 6 shown, the network structure of the ResNet50 feature extraction network can be divided into 5 stages. The first stage consists of the initial convolutional layer 1 and the max pooling layer. The second stage consists of one first bottleneck block and two second bottleneck blocks. The third stage consists of one first bottleneck block and three second bottleneck blocks. The fourth stage consists of one first bottleneck block and five second bottleneck blocks. The fifth stage consists of one first bottleneck block and two second bottleneck blocks. As Figure 7As shown, the first bottleneck block consists of three convolutional layers in series (convolutional layer 2, convolutional layer 3, and convolutional layer 4), another convolutional layer (convolutional layer 5), and a rectified linear unit (RELU). As Figure 8 shown, the second bottleneck block consists of three convolutional layers in series (convolutional layer 6, convolutional layer 7, and convolutional layer 8) and a RELU.

[0125] The convolutional parameters in the above convolutional layers 1 to 8 can be the same or different, and this application does not make any limitations.

[0126] Figures 6 - 8 The ResNet50 feature extraction network shown has a relatively deep network structure, which can learn complex and abstract feature representations. At the same time, due to the existence of residual connections, features can be effectively reused between different layers of the network. This layer-by-layer feature reuse mechanism enables the network to better capture image features at different levels, thereby improving the richness and diversity of features.

[0127] After the feature extraction module extracts the image features of each sub-image, the image features can be passed into the attention module for processing. In the attention module, the attention score of each sub-image can be determined based on the image features of each sub-image, and then based on the attention score and image features of each sub-image, the enhanced feature corresponding to each sub-image can be obtained. Furthermore, the enhanced features corresponding to multiple sub-images can be used as the local features of the image.

[0128] Optionally, the calculation process of the attention score of each sub-image can use the following formula (2), and the calculation process of the enhanced feature corresponding to each sub-image can use the following formula (3).

[0129] Formula (2);

[0130] Formula (3);

[0131] Among them, is a parameter of the person re-identification model, used to increase the dimension; is a parameter of the person re-identification model, used to reduce the dimension; is a rectified linear unit, used to introduce non-linear representation; is a sigmoid activation function, represents the image feature of the i-th sub-image, represents the attention score of the i-th sub-image, represents the enhanced feature corresponding to the i-th sub-image.

[0132] In summary, the embodiments of the present application provide a method for extracting local features. By first segmenting a pedestrian image and then combining an attention mechanism to extract the local features of the segmented sub-images, more local details of the sub-images can be extracted. Based on these local details, different pedestrians can be better distinguished, effectively improving the pedestrian re-identification effect.

[0133] Second, the process of extracting the global features of the image may include: performing global feature extraction processing on the image using a plurality of encoding layers connected in series to obtain the global features of the image.

[0134] Optionally, the above encoding layer is a transformer encoder integrated with wavelet transform.

[0135] See Figure 9 , which shows a schematic structural diagram of an encoding layer. As Figure 9 , the encoding layer is composed of a multi-head self-attention module and a multi-layer perceptron module. Layer normalization is applied before each multi-head self-attention module and multi-layer perceptron module (i.e., normalization module 1 and normalization module 2), and residual connections are applied after each multi-head self-attention module and multi-layer perceptron module.

[0136] Optionally, as Figure 9 , the process of performing global feature extraction processing using any encoding layer may include: obtaining a first feature map input to the encoding layer. Among them, if the encoding layer is the first encoding layer, the first feature map is the initial feature map obtained by initially encoding the sequence of image patches included in the image. If the encoding layer is not the first encoding layer, the first feature map is the second feature map output by the previous encoding layer; processing the first feature map based on the multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to multiple attention heads, performing inverse wavelet transform on the downsampled feature map to obtain a first reconstructed feature map and a second reconstructed feature map. Among them, the downsampled feature map is the feature map obtained during the process of processing the first feature map based on the multi-head attention mechanism integrated with wavelet transform. Based on the query vectors and key-value pairs corresponding to multiple attention heads, the first reconstructed feature map and the second reconstructed feature map, an attention summary map is obtained. Based on the attention summary map and the first feature map, a transition feature map is obtained (corresponding to Figure 9 the normalization module 1, the multi-head self-attention module, and the first residual connection); the second feature map output by the encoding layer is obtained based on the multi-layer perceptron and the transition feature map, where the second feature map output by the last encoding layer is used to obtain the global features of the image (corresponding to Figure 9 the normalization module 2, the multi-layer perceptron module, and the second residual connection).

[0137] Optionally, the process of initially encoding the sequence of image patches included in the image to obtain the initial feature map may include: encoding each image patch in the sequence of image patches included in the image into a vector of a preset length to obtain the vector corresponding to each image patch; performing position encoding on each image patch based on the position of each image patch in the image to obtain the position encoding result corresponding to each image patch; and obtaining the initial feature map based on the vectors and position encoding results respectively corresponding to all the image patches included in the sequence of image patches.

[0138] Optionally, in this embodiment, the sizes of the image patches in the sequence of image patches may be the same, that is, in this embodiment, the image may be divided into multiple image patches of a fixed size to form the sequence of image patches.

[0139] Optionally, a projection method may be adopted to project each image patch in the sequence of image patches into a vector of a preset length.

[0140] Optionally, the calculation formula of the initial feature map is as shown in formula (4) below.

[0141] Formula (4);

[0142] Wherein, represents the initial feature map, represents the vector corresponding to the i-th image patch included in the image, represents the number of image patches included in the image, represents the position encoding results of all the image patches included in the image. Here, contains elements, and the i-th element therein represents the position encoding result corresponding to the i-th image patch.

[0143] As introduced above, in this embodiment, the initial feature map may be input into the first encoding layer as the first feature map to obtain the second feature map output by the first encoding layer; then, the second feature map output by the first encoding layer is input into the second encoding layer as the first feature map to obtain the second feature map output by the second encoding layer; similarly, the second feature map output by the th encoding layer is input into the th encoding layer as the first feature map to obtain the second feature map output by it; and so on, until the second feature map is output by the last encoding layer. Finally, the global feature of the image is obtained based on the second feature map output by the last encoding layer. Wherein, , Indicates the number of encoding layers.

[0144] Taking the process of " using the second feature map output by the -th encoding layer as the first feature map and inputting it into the -th encoding layer to obtain the second feature map output by it" as an example, this process can be obtained using the following formula (5) and formula (6) .

[0145] Formula (5);

[0146] Formula (6);

[0147] Among them, represents the transitional feature map, represents the attention summary map obtained by processing through the multi-head self-attention module MA, represents the normalization process, represents the multi-layer perceptron.

[0148] Optionally, this embodiment can use the following formula (7) to obtain the global feature of the image based on the second feature map output by the last encoding layer .

[0149] Formula (7).

[0150] In a possible implementation, the above process of "processing the first feature map based on the multi-head attention mechanism fused with wavelet transform to obtain query vectors and key-value pairs corresponding to multiple attention heads respectively" may include: normalizing the first feature map to obtain a normalized third feature map; performing a linear transformation on the third feature map using an embedding matrix to obtain a fourth feature map; decomposing the fourth feature map into four wavelet subbands through discrete wavelet transform; concatenating the four wavelet subbands along the channel dimension and then performing a convolution process on the concatenation result to obtain a downsampled feature map; performing multiple groups of linear transformations on the downsampled feature map to obtain query vectors and key-value pairs corresponding to multiple attention heads respectively.

[0151] The above process of "normalizing the first feature map" can be implemented by Figure 9 the normalization module 1.

[0152] Referring to Figure 10 , a schematic structural diagram of a multi-head self-attention module is shown. As Figure 10, the multi-head self-attention module consists of linear layer 1 to linear layer 7, wavelet transform layer 1, wavelet transform layer 2, convolutional layer 9, convolutional layer 10, inverse wavelet transform layer 1, inverse wavelet transform layer 2, and normalization layer 1.

[0153] The above embedding matrix refers to Figure 10 the parameters in linear layer 1 of

[0154] Formula (8);

[0155] Among them, are the parameters in linear layer 1 of the person re-identification model, representing the embedding matrix; represents the third feature map, that is, Figure 10 the input of represents the fourth feature map.

[0156] The above discrete wavelet transform can be implemented by Figure 10 wavelet transform layer 1 and wavelet transform layer 2 of

[0157] Formula (9);

[0158] Formula (10);

[0159] Among them, is the value obtained by wavelet transform, is the discrete scale factor, is the discrete displacement factor, ; is the summation variable, is a wave with oscillation and rapid decay, and satisfies the condition that the integral is 0 in mathematics; is the feature map input to the wavelet transform block.

[0160] The above "performing convolution processing on the splicing result" can be implemented by Figure 10 convolutional layer 9 and convolutional layer 10 of

[0161] Figure 10 Each of linear layer 3, linear layer 4, and linear layer 5 of contains A number of attention heads are used to perform multiple groups of linear transformations on the downsampled feature map to obtain query vectors and key-value pairs corresponding to each of the multiple attention heads.

[0162] Optionally, the process of performing a group of linear transformations on the downsampled feature map to obtain the query vector and key-value pair corresponding to the j-th attention head can be carried out using the following formulas (11) to (13).

[0163] Formula (11);

[0164] Formula (12);

[0165] Formula (13);

[0166] Among them, , and are all parameters in the person re-identification model (i.e., the parameters in linear layer 3, linear layer 4, and linear layer 5), representing the transformation matrix corresponding to the j-th attention head; represents the downsampled feature map, represents the query vector corresponding to the j-th attention head, represents the key corresponding to the j-th attention head, represents the value corresponding to the j-th attention head, , represents the number of attention heads.

[0167] In a possible implementation, the process of "performing inverse wavelet transform on the downsampled feature map to obtain the first reconstructed feature map and the second reconstructed feature map" described above may include: performing two different groups of linear transformations on the downsampled feature map (implemented by the linear layer 2 and linear layer 6 of Figure 10 ) to obtain two different feature maps, and then performing inverse wavelet transforms on the two different feature maps respectively (implemented by the inverse wavelet transform layer 1 and inverse wavelet transform layer 2 of Figure 10 ) to obtain the first reconstructed feature map and the second reconstructed feature map , thereby increasing the receptive field of feature extraction.

[0168] Optionally, the process of "obtaining the attention summary map based on the query vectors and key-value pairs corresponding to the multiple attention heads, the first reconstructed feature map, and the second reconstructed feature map" (which can be implemented by the multiplication unit Figure 10 of , normalization layer 1, and linear layer 7) can be calculated according to the following formula (14).

[0169] Formula (14);

[0170] Among them, represents the feature attention map corresponding to the j-th attention head, and its calculation formula is as follows in formula (15); , and are all parameters in the person re-identification model, representing different transformation matrices; represents the concatenation operation.

[0171] Formula (15);

[0172] Among them, represents the length of, The calculation formula of is as follows in formula (16).

[0173] Formula (16);

[0174] Among them, represents the element in the vector normalized value, is the c-th element in the vector , .

[0175] In summary, in this embodiment, wavelet transform is introduced in the multi-head self-attention module for reversible downsampling, which can avoid the loss of high-frequency information such as texture details in the sampled samples, and at the same time effectively reduce the computational cost. At the same time, the inverse wavelet transform can strengthen the self-attention output by aggregating local context information, further improving the accuracy of the global features.

[0176] After being processed by the multi-head self-attention mechanism and the multi-layer perceptron, the features at each position in the global features can represent the relationship between a local feature in the image and the overall of other regions. Based on this overall relationship, the global features of the image can be captured more accurately from the overall, thereby improving the performance of person re-identification.

[0177] Third, the process of extracting the semantic features of the image may include: performing convolution processing and activation processing on the global features of the image to obtain the fifth feature map corresponding to the image. Among them, the convolution dimension of a convolution function used in the convolution processing includes the number of invariant attributes, and the activation processing is used to enhance the attention to the region where the invariant attributes are located in the image; using a preset attention mechanism to process the global features of the image to obtain the attribute attention map corresponding to the image; and obtaining the semantic features of the image based on the attribute attention map and the fifth feature map corresponding to the image.

[0178] In this embodiment, the process of extracting the semantic features of the image can be implemented by an attribute parsing module. Optionally, the structure of an attribute parsing module is as shown in Figure 11 and this attribute parsing module is composed of convolutional layers 11 to 13, an activation function layer, linear layers 8 to 11, a normalization layer 2, and a pooling layer.

[0179] In this embodiment, the attribute parsing module is located after the last encoding layer mentioned above.

[0180] The above-mentioned "performing convolution processing and activation processing on the global features of the image" can be implemented by convolutional layers 11 to 13 and the activation function layer.

[0181] Optionally, the activation function in the activation function layer is as shown in the following formula (17).

[0182] Formula (17);

[0183] Where , and are scaling factors between 0 and 1, represents the output of the previous convolutional layer of the activation function in the attribute parsing module, represents the activation value, which can be used to obtain the fifth feature map corresponding to the image.

[0184] In this embodiment, the attribute parsing module can also use a preset attention mechanism (i.e., linear layers 8 to 11 and normalization layer 2) to process the global features of the image to obtain the attribute attention map corresponding to the image. Optionally, this process can use the following formula (18).

[0185] Formula (18);

[0186] Where represents the attribute attention map corresponding to the image, represents the preset attention mechanism, respectively represent a set of query vectors, keys, and values obtained based on the global features of the image. Here, the query vectors, keys, and values are different from those mentioned above. Among them, the key is obtained through linear layer 9, the value is obtained through linear layer 8, and the query vector is obtained through linear layer 10.

[0187] The above-mentioned attribute attention map corresponding to the image has a dimension of M×w×h, where M represents the number of invariant attributes, and w and h are the width and height of the attribute attention map corresponding to the image respectively.

[0188] The process of "obtaining the semantic features of the image based on the attribute attention map and the fifth feature map corresponding to the image" can be implemented by Figure 11 a multiplication unit and a pooling layer. Optionally, the process may include: obtaining the attention feature corresponding to each element based on each element included in the attribute attention map corresponding to the image and the fifth feature map, and obtaining the semantic features of the image based on the attention features respectively corresponding to the respective elements .

[0189] Optionally, the process of "obtaining the attention feature corresponding to each element based on each element included in the attribute attention map corresponding to the image and the fifth feature map" (implemented by Figure 11 a multiplication unit ) can be calculated according to the following formula (19).

[0190] Formula (19);

[0191] Wherein, represents the attention feature corresponding to the y-th element included in the attribute attention map corresponding to the image , represents the fifth feature map, represents the y-th element included in the attribute attention map corresponding to the image , .

[0192] Optionally, the process of "obtaining the semantic features of the image based on the attention features respectively corresponding to the respective elements" (implemented by Figure 11 a pooling layer) may include: performing generalized mean pooling on the attention features respectively corresponding to the respective elements to obtain the semantic features of the image.

[0193] In summary, in this embodiment, the activation function provided by the attribute parsing module enhances the attention of the person re-identification model to the regions where the relevant invariant attributes are located, while reducing potential biases, making the obtained fifth feature map more accurate, and thus improving the accuracy of the semantic features of the image.

[0194] In some other embodiments of the present application, the process of fusing the local features, global features, and semantic features of any image in step S402 is introduced.

[0195] Optionally, in this embodiment, weights may be pre-configured for the local features, global features, and semantic features of the image, and then the local features, global features, and semantic features of the image may be weighted and fused through the pre-configured weights.

[0196] Optionally, the above-mentioned pre-configured weights can be configured based on the importance degrees of local features, global features, and semantic features respectively. For example, for the above-mentioned local features including head features , upper body features and lower body features as an example, due to the differences in height, pitch angle, etc. between the UAV perspective and the ground monitoring perspective, the same body parts of pedestrians in different perspectives may have different importance. For example, the UAV flies at a relatively high altitude, so the features of the head in the pedestrian images captured by it are more obvious than those of the lower body, and a larger weight should be pre-configured.

[0197] Optionally, the above-mentioned weighted fusion process can adopt the following formula (20).

[0198] Formula (20);

[0199] Wherein, represents the fused feature, , , , and respectively represent the pre-configured weights of the global feature , head feature , upper body feature , lower body feature and semantic feature respectively.

[0200] In this embodiment, contains a comprehensive and detailed representation of pedestrians, provides a reliable distance calculation basis for cross-view pedestrian re-identification based on the UAV perspective and the ground monitoring perspective, and can perform pedestrian re-identification more accurately based on this fused feature.

[0201] The following embodiment introduces the loss function of the above-mentioned pedestrian re-identification model.

[0202] Optionally, the loss function of the pedestrian re-identification model is a joint loss function composed of a triplet loss function, a metric distillation loss function, and a cross-entropy loss function.

[0203] The following formula (21) is the calculation formula of the metric distillation loss function.

[0204] Formula (21);

[0205] Wherein, represents the image samples and image samples The corresponding metric distillation loss, where M is the number of invariant attributes, denotes the image sample (i.e., the training ground surveillance image) and the image sample (i.e., the training drone image); denotes the distance between the image sample and the image sample under the influence of the y-th invariant attribute, which is calculated from the respective attention features of the two image samples.

[0206] The following formula (22) is the calculation formula of the triplet loss function.

[0207] Formula (22);

[0208] where, denotes the triplet loss corresponding to the image sample contained in the training data, denotes the predefined distance.

[0209] This triplet loss function aims to ensure that for any given image sample , the distance to the negative sample is greater than the distance to the positive sample plus the predefined distance .

[0210] The following formula (23) is the calculation formula of the cross-entropy loss function.

[0211] Formula (23);

[0212] where, denotes the number of image samples contained in the training data, denotes the predicted probability of the person re-identification model for the i-th image sample, denotes the one-hot encoded true label corresponding to the i-th image sample, denotes the cross-entropy loss corresponding to the i-th image sample contained in the training data.

[0213] Optionally, this embodiment can calculate the total loss of the model using the following formula (24).

[0214] Formula (24);

[0215] where, and are the parameters of the person re-identification model, and the above represents the sum of the metric distillation losses corresponding to all image pairs (an image pair includes two image samples in the training data) included in the training data, represents the sum of the triplet losses corresponding to all image samples included in the training data, represents the sum of the cross-entropy losses corresponding to all image samples included in the training data, represents the total loss of the model.

[0216] In summary, in this embodiment, the metric distillation loss, the triplet loss, and the cross-entropy loss are combined, so that the person re-identification model can be optimized with the goal of making the fusion feature distances of the same person under different perspectives as equal as possible, effectively improving the performance of the person re-identification model for person re-identification.

[0217] The above introduces a person re-identification method provided by an embodiment of the present application. Next, a device for executing the above person re-identification method will be introduced.

[0218] Please refer to Figure 12 , Figure 12 , which is a schematic structural diagram of a person re-identification device provided by an embodiment of the present application. As Figure 12 shown, the device may include:

[0219] An image acquisition module 501, configured to acquire a ground surveillance image of a target person and a set of drone images to be recognized, where the set of drone images to be recognized includes multiple drone images;

[0220] A person re-identification module 502, configured to respectively extract local features, global features, and semantic features of the ground surveillance image and each drone image by using a person re-identification model, so as to determine a drone image of the target person from the set of drone images to be recognized based on the extracted features;

[0221] Among them, the person re-identification model uses the training ground surveillance images and the training set of drone images with labeled training person labels as training data. Each drone image included in the training set of drone images has a perspective difference and a scale difference from the training ground surveillance images. The global features represent the overall appearance information of the person and the information of the person and the surrounding environment. The local features represent the detailed information of different regions of the person's body. The semantic features represent the information of the invariant attributes on the person.

[0222] In a possible implementation, when the above person re-identification module extracts local features of any one of the ground surveillance image and each drone image, it may specifically be used for:

[0223] Segment the image according to body parts to obtain a plurality of segmented sub-images;

[0224] Extract the image features of each of the multiple sub-images using a feature extraction network;

[0225] Perform attention-based feature enhancement processing on the image features of each of the multiple sub-images respectively to obtain enhanced features corresponding to each of the multiple sub-images;

[0226] Use the enhanced features corresponding to each of the multiple sub-images as local features of the image.

[0227] In a possible implementation, when the above pedestrian re-identification module extracts the global feature of any one of the ground surveillance image and each UAV image, it can specifically be used for:

[0228] Perform global feature extraction processing on the image using multiple cascaded encoding layers to obtain the global feature of the image;

[0229] Among them, the process of performing global feature extraction processing using any one encoding layer includes:

[0230] Obtain the first feature map input to the encoding layer. Among them, if the encoding layer is the first encoding layer, the first feature map is the initial feature map obtained by performing preliminary encoding on the sequence of image patches included in the image. If the encoding layer is not the first encoding layer, the first feature map is the second feature map output by the previous encoding layer;

[0231] Process the first feature map based on the multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to each of the multiple attention heads;

[0232] Perform inverse wavelet transform on the downsampled feature map to obtain the first reconstructed feature map and the second reconstructed feature map, where the downsampled feature map is the feature map obtained during the process of processing the first feature map based on the multi-head attention mechanism integrated with wavelet transform;

[0233] Based on the query vectors and key-value pairs corresponding to each of the multiple attention heads, the first reconstructed feature map and the second reconstructed feature map, obtain an attention summary map;

[0234] Based on the attention summary map and the first feature map, obtain a transition feature map;

[0235] Based on the multi-layer perceptron and the transition feature map, obtain the second feature map output by the encoding layer, where the second feature map output by the last encoding layer is used to obtain the global feature of the image.

[0236] In a possible implementation, when the above pedestrian re-identification module processes the first feature map based on the multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to each of the multiple attention heads, it can specifically be used for:

[0237] Normalize the first feature map to obtain a normalized third feature map;

[0238] Perform a linear transformation on the third feature map using an embedding matrix to obtain a fourth feature map;

[0239] Decompose the fourth feature map into four wavelet subbands through discrete wavelet transform;

[0240] Concatenate the four wavelet subbands along the channel dimension and then perform convolution processing on the concatenated result to obtain a downsampled feature map;

[0241] Perform multiple groups of linear transformations on the downsampled feature map to obtain query vectors and key-value pairs corresponding to multiple attention heads.

[0242] In a possible implementation, when the above pedestrian re-identification module performs preliminary encoding, it can specifically be used for:

[0243] Encode each image patch in the sequence of image patches included in the image into a vector of a preset length to obtain a vector corresponding to each image patch;

[0244] Perform position encoding on each image patch based on the position of the image patch in the image to obtain a position encoding result corresponding to each image patch;

[0245] Based on the vectors and position encoding results respectively corresponding to all the image patches included in the image patch sequence, obtain an initial feature map.

[0246] In a possible implementation, when the above pedestrian re-identification module extracts semantic features from any one of the ground surveillance image and each drone image, it can specifically be used for:

[0247] Perform convolution processing and activation processing on the global features of the image to obtain a fifth feature map corresponding to the image, where the convolution dimension of a convolution function used in the convolution processing includes the number of invariant attributes, and the activation processing is used to enhance the attention to the region where the invariant attributes are located in the image;

[0248] Process the global features of the image using a preset attention mechanism to obtain an attribute attention map corresponding to the image;

[0249] Based on the attribute attention map and the fifth feature map corresponding to the image, obtain the semantic features of the image.

[0250] In a possible implementation, the loss function of the pedestrian re-identification model in the above pedestrian re-identification module is a joint loss function composed of a triplet loss function, a metric distillation loss function, and a cross-entropy loss function.

[0251] The pedestrian re-identification device provided in this application corresponds to the aforementioned pedestrian re-identification method. For details, please refer to the previous introduction and will not be elaborated here.

[0252] An electronic device is also provided in an embodiment of this application. Refer to Figure 13 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of this application. The electronic device in the embodiment of this application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and so on. Figure 13 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiment of this application.

[0253] As Figure 13 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 608 into the random access memory (RAM) 603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.

[0254] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a memory card, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 13 the electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.

[0255] A computer program product is also provided in an embodiment of this application, including computer-readable instructions. When the computer-readable instructions run on the electronic device, the electronic device is enabled to implement any one of the pedestrian re-identification methods provided in the embodiment of this application.

[0256] A computer-readable storage medium is also provided in an embodiment of this application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by the electronic device, the electronic device can be enabled to implement any one of the pedestrian re-identification methods provided in the embodiment of this application.

[0257] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.

[0258] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0259] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0260] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. A pedestrian re-identification method, characterized in that: include: Acquire a ground surveillance image of a target pedestrian and a set of unmanned aerial vehicle images to be identified, wherein the set of unmanned aerial vehicle images to be identified includes multiple unmanned aerial vehicle images; Using a pedestrian re-identification model, respectively extract local features, global features, and semantic features of the ground surveillance image and each of the drone images, so as to determine the drone image of the target pedestrian from the set of drone images to be identified based on the extracted features; The pedestrian re-identification model uses training ground surveillance images and training drone image sets marked with training pedestrian labels as training data. Each drone image included in the training drone image set has a perspective difference and a scale difference from the training ground surveillance image. The global features represent the overall appearance information of the pedestrian and the information about the pedestrian and the surrounding environment. The local features represent the detailed information of different areas of the pedestrian's body. The semantic features represent the information of the invariant attributes of the pedestrian. For any one of the ground monitoring image and each of the drone images, a process of extracting the global features of the image includes: A plurality of coding layers connected in series are used to perform global feature extraction processing on the image to obtain the global features of the image; The process of using any one of the coding layers to perform global feature extraction processing includes: Obtaining a first feature map input to the encoding layer, wherein if the encoding layer is the first encoding layer, the first feature map is an initial feature map obtained by preliminarily encoding a sequence of image blocks contained in the image; if the encoding layer is not the first encoding layer, the first feature map is a second feature map output by a previous encoding layer; Processing the first feature map based on a multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to multiple attention heads respectively; Performing an inverse wavelet transform based on the downsampled feature map to obtain a first reconstructed feature map and a second reconstructed feature map, wherein the downsampled feature map is a feature map obtained in the process of processing the first feature map based on the multi-head attention mechanism fused with wavelet transform; Obtaining an attention summary graph based on the query vectors and key-value pairs respectively corresponding to the multiple attention heads, the first reconstructed feature graph, and the second reconstructed feature graph; Obtaining a transition feature map based on the attention summary map and the first feature map; A second feature map output by the encoding layer is obtained based on the multi-layer perceptron and the transition feature map, wherein the second feature map output by the last encoding layer is used to obtain the global feature of the image.

2. The pedestrian re-identification method according to claim 1, characterized in that: For any one of the ground monitoring image and each of the drone images, a process of extracting local features of the image includes: Segmenting the image according to body parts to obtain multiple segmented sub-images; Using a feature extraction network to extract image features of each of the plurality of sub-images; Performing attention-based feature enhancement processing on the image features of each of the plurality of sub-images to obtain enhanced features corresponding to the plurality of sub-images; The enhanced features respectively corresponding to the multiple sub-images are used as local features of the image.

3. The pedestrian re-identification method according to claim 1, characterized in that: The multi-head attention mechanism based on wavelet transform is used to process the first feature map to obtain query vectors and key-value pairs corresponding to multiple attention heads, including: Normalizing the first feature map to obtain a normalized third feature map; Performing a linear transformation on the third feature map using an embedding matrix to obtain a fourth feature map; Decomposing the fourth feature map into four wavelet sub-bands by discrete wavelet transform; The four wavelet sub-bands are spliced ​​together along the channel dimension, and then a convolution process is performed on the splicing result to obtain the down-sampled feature map; Perform multiple sets of linear transformations on the downsampled feature map to obtain query vectors and key-value pairs corresponding to the multiple attention heads.

4. The pedestrian re-identification method according to claim 1, characterized in that: The preliminary coding process includes: Encoding each image block in the image block sequence contained in the image into a vector of a preset length, and obtaining a vector corresponding to each image block; Based on the position of each image block in the image, position encoding is performed on each image block to obtain a position encoding result corresponding to each image block; The initial feature map is obtained based on the vectors and position encoding results corresponding to all the image blocks included in the image block sequence.

5. The pedestrian re-identification method according to claim 1, characterized in that: For any one of the ground monitoring image and each of the drone images, a process of extracting the semantic features of the image includes: Performing convolution processing and activation processing on the global features of the image to obtain a fifth feature map corresponding to the image, wherein a convolution dimension of a convolution function used in the convolution processing includes the number of the invariant attributes, and the activation processing is used to enhance the attention to the area where the invariant attributes are located in the image; The global features of the image are processed using a preset attention mechanism to obtain an attribute attention map corresponding to the image; Based on the attribute attention map and the fifth feature map corresponding to the image, the semantic features of the image are obtained.

6. The pedestrian re-identification method according to any one of claims 1 to 5, characterized in that: The loss function of the pedestrian re-identification model is a joint loss function composed of a triplet loss function, a metric distillation loss function and a cross entropy loss function.

7. A pedestrian re-identification device, characterized in that: include: An image acquisition module, used to acquire a ground surveillance image of a target pedestrian and a set of unmanned aerial vehicle images to be identified, wherein the set of unmanned aerial vehicle images to be identified includes multiple unmanned aerial vehicle images; A pedestrian re-identification module, configured to extract local features, global features, and semantic features of the ground surveillance image and each of the drone images respectively by using a pedestrian re-identification model, so as to determine the drone image of the target pedestrian from the set of drone images to be identified based on the extracted features; The pedestrian re-identification model uses training ground surveillance images and training drone image sets marked with training pedestrian labels as training data. Each drone image included in the training drone image set has a perspective difference and a scale difference from the training ground surveillance image. The global features represent the overall appearance information of the pedestrian and the information about the pedestrian and the surrounding environment. The local features represent the detailed information of different areas of the pedestrian's body. The semantic features represent the information of the invariant attributes of the pedestrian. For any one of the ground monitoring image and each of the drone images, a process of extracting the global features of the image includes: A plurality of coding layers connected in series are used to perform global feature extraction processing on the image to obtain the global features of the image; The process of using any one of the coding layers to perform global feature extraction processing includes: Obtaining a first feature map input to the encoding layer, wherein if the encoding layer is the first encoding layer, the first feature map is an initial feature map obtained by preliminarily encoding a sequence of image blocks contained in the image; if the encoding layer is not the first encoding layer, the first feature map is a second feature map output by a previous encoding layer; Processing the first feature map based on a multi-head attention mechanism integrated with wavelet transform to obtain query vectors and key-value pairs corresponding to multiple attention heads respectively; Performing an inverse wavelet transform based on the downsampled feature map to obtain a first reconstructed feature map and a second reconstructed feature map, wherein the downsampled feature map is a feature map obtained in the process of processing the first feature map based on the multi-head attention mechanism fused with wavelet transform; Obtaining an attention summary graph based on the query vectors and key-value pairs respectively corresponding to the multiple attention heads, the first reconstructed feature graph, and the second reconstructed feature graph; Obtaining a transition feature map based on the attention summary map and the first feature map; A second feature map output by the encoding layer is obtained based on the multi-layer perceptron and the transition feature map, wherein the second feature map output by the last encoding layer is used to obtain the global feature of the image.

8. An electronic device, characterized in that: The method comprises at least one processor and a memory connected to the processor, wherein: The memory is used to store computer programs; The processor is used to execute the computer program so that the electronic device can implement the pedestrian re-identification method as described in any one of claims 1 to 6.

9. A computer storage medium, characterized in that: The storage medium carries one or more computer programs, and when the one or more computer programs are executed by an electronic device, the electronic device can implement the pedestrian re-identification method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Double-branch adaptive coding and decoding colorectal polyp segmentation method

    CN117765012A

  • Air-ground cross-platform target re-identification method based on semantic alignment and prompt learning

    CN119007241A