Local feature optimization-based visible light and infrared pedestrian re-identification method

By constructing local feature extraction, alignment and collaborative learning modules in pedestrian recognition, the limitations of the failure to fully utilize image modal differences in the prior art are solved, and more efficient cross-modal pedestrian feature identification is achieved.

WO2025091620A1PCT designated stage expired Publication Date: 2025-05-08SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI

Patent Information

Application Number
PCT/CN2023/137192
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2023-12-07
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The prior art fails to fully consider the difference in image modality when extracting local features in visible light and infrared pedestrian images, resulting in the limitation of discrimination of cross-modal local features.

Method used

Using a method based on local feature optimization, the global feature extraction module, local feature extraction module, local feature alignment module and feature collaborative learning module are used to adaptively extract and align cross-modal local features, and enhance the discriminantity of local features through feature collaborative learning.

Benefits of technology

It significantly improves the optimization effect of visible light and infrared cross-modal local features, enhances the discrimination of pedestrian identity characteristics, and improves the accuracy of pedestrian re-identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023137192_08052025_PF_FP_ABST
    Figure CN2023137192_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a local feature optimization-based visible light and infrared pedestrian re-identification method. The method comprises: collecting a target area image, the target area image comprising a visible light-mode image and an infrared-mode image; and inputting the target area image into a pedestrian identification model to obtain a pedestrian identification result, wherein the pedestrian identification model comprises a local feature extraction module, a local feature alignment module and a feature collaborative learning module, the local feature extraction module extracts local features of the visible light mode and infrared mode, the local feature alignment module groups and aligns cross-modal local features, so as to establish cross-modal connections, and the feature collaborative learning module enhances each local feature and performs collaborative optimization on the global features and the local features. The present invention can adaptively mine cross-modal local features, directly establish rich alignment relationships for the cross-modal local features, and embed context relationships of global features into the local features, thus remarkably improving the accuracy of pedestrian identification.
Need to check novelty before this filing date? Find Prior Art

Description

A visible light and infrared pedestrian re-identification method based on local feature optimization Technical Field

[0001] The present invention relates to the field of computer vision technology, and more particularly to a visible light and infrared pedestrian re-identification method based on local feature optimization. Background Art

[0002] Person re-identification (PRED) is a technology that automatically analyzes video to retrieve pedestrians with the same identity from different cameras. Significant modal differences exist between visible light and infrared images in both daytime and nighttime scenarios, complicating PRED. Existing technologies still face significant challenges in extracting local information from visible and infrared modalities to identify pedestrians. This includes identifying clothing patterns, clothing length, shoe type, backpack type, and other subtle and distinctive local features.

[0003] Typically, visible light images of pedestrians can be used to identify individuals based on clothing color information, but they also have numerous untapped local features. In contrast, infrared images lack distinct features and rely on detailed features for accurate identification. Intuitively, the discriminability of infrared features is often less than that of visible light. Therefore, without being able to detect these subtle differences in pedestrian features, directly matching pedestrians from different modalities presents challenges. Overall, local features offer significant potential for improving the discriminability of visible and infrared features, playing a crucial role in cross-modal person re-identification.

[0004] Currently, the methods for cross-modal person re-identification using local features can be mainly divided into two categories. The first category mainly uses a uniform partitioning method to extract local features. For example, given a pedestrian image, the model outputs several local features corresponding to the pedestrian image. These local features are obtained by a fixed partitioning of the global features, and then different local features are trained using a loss function. This type of solution obtains local features through a rough partitioning method and then performs post-optimization, emphasizing the consistency of the size of each local feature space. However, due to insufficient consideration of modal changes, this uniform partitioning method makes it difficult to establish connections between different cross-modal local features, thereby limiting the cross-modal discriminability of local features.

[0005] The second type of solution mainly uses the auxiliary information of pedestrians to guide the model to extract the local features of pedestrians. For example, unsupervised migration is performed using a posture estimation model trained on other datasets to obtain the positions of local parts such as the skeleton key points of pedestrians in the cross-modal pedestrian re-identification dataset, and then the local features of pedestrians are extracted based on these positions. This type of solution usually locates the local parts of pedestrians, but requires the help of additional labeled prior information, such as pedestrian key points, pedestrian attributes, or manually parsed information. The acquisition of this information requires relying on human posture estimation datasets and complex posture estimation models. In addition, there are certain deviations between posture estimation and cross-modal pedestrian re-identification datasets. These deviations will hinder the ideal semantic segmentation of pedestrian images.

[0006] In summary, existing schemes do not fully consider the impact of image modality differences in the process of extracting local features, and thus the optimization of local discriminant features of different modalities is subject to certain limitations.

[0007] Summary of the Invention

[0008] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide a method for visible light and infrared pedestrian re-identification based on local feature optimization. The method comprises the following steps:

[0009] Acquiring a target area image, wherein the target area image includes a visible light modality image and an infrared modality image;

[0010] Inputting the target area image into a trained pedestrian recognition model to obtain a pedestrian identification result;

[0011] Among them, the pedestrian recognition model includes a global feature extraction module, a local feature extraction module, a local feature alignment module, a feature collaborative learning module and a classification module. The global feature extraction module is used to extract visible light modal global features and infrared modal global features; the local feature extraction module is used to extract corresponding visible light modal local features and infrared modal local features from the visible light modal global features and infrared modal global features respectively; the local feature alignment module is used to group the visible light modal local features and infrared modal local features to obtain multiple cross-modal combinations, and then align each group of cross-modal local features and distinguish different groups of cross-modal local features; the feature collaborative learning module is used to embed the visible light modal global features and infrared modal global features into the corresponding local features to obtain enhanced local features, and collaboratively optimize the enhanced local features with the global features; the output features of the feature collaborative learning module are passed to the classification module to classify the pedestrian identity.

[0012] Compared with the existing technology, the advantage of the present invention is that the provided visible light and infrared pedestrian re-identification method based on local feature optimization does not adopt a fixed division method when extracting local features, nor does it rely on additional human auxiliary information. Instead, it attempts to guide the network to adaptively select some local features with pedestrian discrimination value, and then align these local features across modalities. The entire process ensures the full mining of fine-grained information while also continuously prompting local features to adapt to modal changes. The present invention enhances the discriminability of cross-modal identity features and can better improve the optimization effect of visible light and infrared cross-modal local features.

[0013] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0015] FIG1 is a flowchart of a method for visible light and infrared pedestrian re-identification based on local feature optimization according to an embodiment of the present invention;

[0016] FIG2 is an architecture diagram of a pedestrian recognition model according to an embodiment of the present invention;

[0017] FIG3 is a visualization diagram of a local feature heat map according to an embodiment of the present invention. DETAILED DESCRIPTION

[0018] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention.

[0019] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0020] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0021] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0022] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0023] In order to mine local information in pedestrian images and further improve the discriminability of cross-modal pedestrian features, the present invention proposes a visible light and infrared pedestrian re-identification method based on local feature optimization. This method mainly focuses on three aspects: extraction, alignment, and enhancement of local features. First, a local feature extraction module is constructed to extract local features of different modalities. Then, in order to make full use of the rich detail information in the features to establish cross-modal connections, a local feature alignment module is constructed to group and align cross-modal local features. In addition, a feature collaborative learning module is constructed to enhance each local feature, while also promoting the joint optimization of global and local features.

[0024] Specifically, referring to FIG1 , the provided visible light and infrared pedestrian re-identification method based on local feature optimization includes the following steps:

[0025] Step S110 , constructing a pedestrian recognition model, which generally includes a global feature extraction module, a local feature extraction module, a local feature alignment module, a feature collaborative learning module and a classification module.

[0026] As shown in FIG2 , the pedestrian recognition model includes a global feature extraction module 10 , a local feature extraction module 20 , a local feature alignment module 30 , a feature collaborative learning module 40 , and a classification module 50 .

[0027] The global feature extraction module 10 is used to extract global features from both visible light and infrared image modalities, obtaining visible light modality global features (or simply visible light global features) and infrared modality global features (or simply infrared global features). The global feature extraction module 10 can be implemented using existing deep neural networks, such as convolutional neural networks or recurrent neural networks.

[0028] The local feature extraction module 20 extracts adaptive modality-specific local features, or referred to as visible light modality local features and infrared modality local features, from the visible light modality global features and the infrared modality global features, respectively.

[0029] The local feature alignment module 30 is used to group and align the local features of the two modalities.

[0030] The feature collaborative learning module 40 is used to embed the context information of the global feature into each local feature, further enhance the discriminability of the local feature, and collaboratively optimize the enhanced local feature with the global feature.

[0031] The features output by the feature collaborative learning module 50 are input to the classification module 50 to classify the pedestrian features. The classification module 50 can be implemented using a fully connected layer.

[0032] In one embodiment, the local feature extraction module mainly consists of two parameter-learnable global feature segmenters G v and G t Composition, the global characteristics of visible light and infrared global features Send to G v and G t After that, we get the S-block local attention map of the visible light modality And the S-block local attention map of the infrared modality Then, and Respectively with local attention map and Multiply to obtain specific local features of visible light and infrared specific local features The specific process can be expressed as:

[0033] Among them, σ is the sigmoid activation function, represents a 1×1 convolution operation, A v Focusing on the different details of the global features of visible light, A t The focus is on different detail information in the infrared global features.

[0034] In order to ensure that the modality-specific local features can capture different fine-grained information, in one embodiment, a feature segmentation loss function is designed to force each local attention map in the same modality to focus on different parts of the pedestrian. For example, the feature segmentation loss The calculation expression is:

[0035] Among them, S represents the number of local attention maps, that is, the number of local features. It means the i-th row and j-th column In the same modality, the feature segmentation loss guides each local attention map to extract fine-grained information from different parts of the pedestrian by minimizing the overlapping area between different local attention maps.

[0036] In one embodiment, the specific process of the local feature alignment module is to first group the local features of the visible light modality and the infrared modality to obtain s combinations, and then align each group of cross-modal local features, while distinguishing different groups of cross-modal local features. In this way, each group of cross-modal local features relies on different detailed information during the alignment process. The group alignment process can guide the model to discover more local information that is suitable for alignment and non-repetitive, such as backpacks, clothing patterns, pants length, etc.

[0037] Each set of cross-modal local features is fed into different convolution modules conv i =[conv1,conv2,…,conv s ], and then obtain the visible light alignment local features Align local features with infrared The calculation expression is as follows:

[0038] Among them, all convolution modules conv are composed of convolution blocks with a convolution kernel size of 1 and a nonlinear activation function ReLU. The structure remains consistent, but the parameters are independent of each other.

[0039] In order to use these convolutional modules to align cross-modal local features of the same group and distinguish cross-modal local features of different groups, in one embodiment, a local group alignment loss is proposed. The calculation expression is as follows:

[0040] Among them, i and j represent the group numbers of different cross-modal local features, there are s groups in total, and cos(x,y) represents the cosine similarity between x and y, which is defined as In addition, the method uses a monotonically increasing ln(1+exp(·)) function to maintain training stability by avoiding negative loss values. Designing a local group alignment loss can continuously improve the similarity of cross-modal local features within the same group and reduce the similarity of cross-modal local features between different groups, achieving the goal of aligning cross-modal local features within a group and distinguishing between groups.

[0041] In one embodiment, the specific process of the feature collaborative learning module is as follows: first, information interaction between global features and local features is established through a designed feature aggregation method, and each local feature is enhanced separately using the global feature; then, the global feature and the enhanced local feature are spliced ​​together to obtain the final optimized feature.

[0042] To explain the feature aggregation method mentioned above, the global features in the visible light modality are used below. and the i-th local feature Let's introduce the specific process of feature aggregation.

[0043] First, the global features and local features Perform global average pooling and map it to the same feature space through a Linear layer composed of a fully connected layer to obtain and Expressed as:

[0044] where g(·) represents the GAP (global average pooling) operation.

[0045] Then, and Perform dot product operation to obtain the aggregation relationship graph E, which is expressed as:

[0046] Among them, the softmax function is defined as The weight of the aggregation relationship graph is used to measure the relationship between global features and local features.

[0047] Next, in order to integrate global and local information, the response features are calculated Afterwards, and Perform spatial fusion to obtain enhanced local features Expressed as:

[0048] Among them, ω i Represents a learnable adaptive weight.

[0049] In summary, the feature collaborative learning module first embeds global information into each local feature through feature aggregation, and uses the relationship between global features and local features to enhance the robustness of local features. Then, the enhanced local features and global features are spliced ​​in the channel dimension to promote mutual optimization between them.

[0050] Step S120 , minimizing the overall loss function as the optimization goal, and using the known data set to train the pedestrian recognition model.

[0051] In this step, a pedestrian recognition model is trained using a known dataset to optimize the model's learnable parameters until a set optimization objective is achieved. For example, the optimization objective can be minimizing the overall loss function. Each sample in the dataset reflects the correspondence between a visible light or infrared modality image and a pedestrian identity.

[0052] Specifically, the overall training process is as follows: pedestrian images in both visible and infrared modalities are read from the dataset. In each training batch, each pedestrian identity has an equal number of visible and infrared images. The training images are fed into a preselected deep neural network to extract global features from both modalities. The local feature extraction module extracts adaptive modality-specific local features from both visible and infrared global features. The local feature alignment module groups and aligns the local features of the two modalities. The feature collaborative learning module embeds contextual information from the global features into each local feature to further enhance the discriminability of the local features, and then collaboratively optimizes the enhanced local features with the global features. The features output by the feature collaborative learning module are fed into the classification module to classify pedestrian features.

[0053] In one embodiment, the overall loss function for training a pedestrian recognition model is expressed as:

[0054] Among them, α and β are hyperparameters, and identity loss and triplet loss Mainly used for classification, the calculation expressions are:

[0055] Where N represents the number of pedestrian images of the same modality in each training batch, m represents the number of pedestrian categories in the cross-modal training dataset, subscript i represents the index of the pedestrian image in the current batch, and subscript j represents the category of the pedestrian. represents the parameters of the fully connected layer used for classification, and f represents the pedestrian features extracted from the pedestrian image through the model.

[0056] Where y represents the category corresponding to the feature of the input sample, and ρ1 is a predefined interval value. The triplet loss mainly takes a triplet (anchor sample a, positive sample p and negative sample n) as input, where p and a have the same label, while n and a have different labels. After the model extracts features from the input sample, it obtains the positive sample feature pair {f s ,f p} and negative sample feature pairs {f a ,f n}, continuously driving the distance between the positive sample feature pairs to be less than a set value than the distance between the negative sample feature pairs, thereby achieving feature clustering.

[0057] It should be noted that the above-mentioned identity loss and triplet loss can be general losses, or other classification losses can be used.

[0058] Step S130 : For the collected visible light modality image and infrared modality image, the identity of the pedestrian is recognized in real time using the trained pedestrian recognition model.

[0059] After the model training is completed, it can be used for actual pedestrian re-identification. For example, the real-time collected visible light modal images and infrared modal images can be input into the trained pedestrian recognition model to identify the identity of the pedestrian.

[0060] It should be understood that the training process of the pedestrian recognition model described in the present invention can be performed offline on a server or in the cloud. By embedding the trained model into an electronic device, real-time pedestrian identification can be achieved. The electronic device can be a terminal device or a server. Terminal devices include any terminal device, such as monitoring equipment, mobile phones, tablet computers, personal digital assistants (PDAs), in-vehicle computers, and smart wearable devices (smart watches, virtual reality glasses, virtual reality helmets, etc.). Servers include, but are not limited to, application servers or web servers, and can be standalone servers, cluster servers, or cloud servers.

[0061] In order to further verify the effect of the present invention, an ablation experiment was conducted and the experimental results were analyzed using a visualization method.

[0062] (1) Ablation experiment

[0063] Experiments were conducted on two mainstream visible light and infrared cross-modal pedestrian re-identification datasets. The information of these two datasets is as follows.

[0064] SYSU-MM01: This dataset contains 491 pedestrian identities, 287,628 visible light images, and 15,792 infrared images. The training set contains 395 pedestrian identities, and the test set contains 96 pedestrian identities. The entire dataset is collected from four visible light cameras and two near-infrared cameras.

[0065] RegDB: This dataset contains 412 pedestrian identities, with 4120 visible light images and 4120 infrared images each. Each pedestrian with the same identity has 10 visible light images and 10 infrared images. The training and test sets each contain images of 206 pedestrian identities. The entire dataset is collected from one visible light camera and one thermal infrared camera.

[0066] Ablation experiments measure model performance using three evaluation metrics: Rank-1, Rank-10, and mean average precision (mAP). These metrics are numbers between 0 and 1, with larger values ​​representing higher person re-ID accuracy. Tables 1 and 2 illustrate experimental results on the SYSU-MM01 and RegDB datasets, using a baseline network based on ResNet50 as the backbone. The local feature extraction module (marked with the letter A), local feature alignment module (marked with the letter B), and feature collaborative learning module (marked with the letter C) included in this invention were applied to the baseline network.

[0067] Table 1 Ablation experiment results on the SYSU-MM01 dataset (%)

[0068] Table 2 Ablation experiment results on RegDB dataset (%)

[0069] It can be seen from Table 1 and Table 2 that on the same training set, compared with the baseline network, the present invention can significantly improve the values ​​of the three indicators of Rank-1, Rank-10 and average accuracy mAP of the model in the test set.

[0070] (2) Heat map visualization

[0071] In order to prove that the present invention can mine fine-grained information in pedestrian features and enhance the discriminability of pedestrian features, the Grad-CAM method is used in the experiment to visualize the local features of the two modalities and highlight the key areas of focus on the corresponding pedestrian images.

[0072] In the experiment, the number of local features is set to 5. As shown in Figure 3, each row represents the input image and the corresponding 5 local feature heat maps, and the i-th column (i = 1, 2, 3, 4, 5) represents the visible light local feature. or infrared local features Figure 3(a) and Figure 3(b) show the visualization results of the same pedestrian, with visible light and infrared images, respectively. As can be seen in Figure 3, each local feature focuses on corresponding details, and the focus area of ​​each heatmap is somewhat different from that of other heatmaps. This indicates that the feature segmentation loss can guide the local feature extraction module to extract local features that focus on different parts of the body.

[0073] In summary, the visible light and infrared pedestrian re-identification method based on local feature optimization provided by the present invention obtains different local features in pedestrian images, guides the model to pay attention to subtle identity clues in cross-modal features, and improves the accuracy of local feature extraction. In addition, the present invention does not use a fixed horizontal division method to extract local features, nor does it require any human auxiliary information. Instead, it guides the model to mine adaptive cross-modal local features and directly establishes rich alignment relationships for cross-modal local features. At the same time, it also considers the relationship between global features and local features, embeds the contextual relationship of global features into local features, and significantly improves the recognition performance of the model.

[0074] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0075] Computer-readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device.Computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof.More specific examples (non-exhaustive list) of computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove having instructions stored thereon, and any suitable combination thereof.Computer-readable storage medium used herein is not interpreted as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagated by waveguides or other transmission media (for example, light pulses by fiber optic cables), or electrical signals transmitted by wires.

[0076] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0077] The computer program instructions for performing the operation of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, and conventional procedural programming languages ​​such as "C" language or similar programming languages. The computer readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), is personalized by utilizing the state information of the computer readable program instructions, and the electronic circuit can execute the computer readable program instructions, thereby realizing various aspects of the present invention.

[0078] Various aspects of the present invention are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0079] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0080] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0081] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.

[0082] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A visible light and infrared pedestrian re-identification method based on local feature optimization, comprising the following steps: Acquiring a target area image, wherein the target area image includes a visible light modality image and an infrared modality image; Inputting the target area image into a trained pedestrian recognition model to obtain a pedestrian identification result; Among them, the pedestrian recognition model includes a global feature extraction module, a local feature extraction module, a local feature alignment module, a feature collaborative learning module and a classification module. The global feature extraction module is used to extract visible light modal global features and infrared modal global features; the local feature extraction module is used to extract corresponding visible light modal local features and infrared modal local features from the visible light modal global features and the infrared modal global features respectively; the local feature alignment module is used to group the visible light modal local features and the infrared modal local features to obtain multiple cross-modal combinations, and then align each group of cross-modal local features, and distinguish different groups of cross-modal local features; the feature collaborative learning module is used to embed the visible light modal global features and the infrared modal global features into the corresponding local features to obtain enhanced local features, and collaboratively optimize the enhanced local features with the global features; the output features of the feature collaborative learning module are passed to the classification module to classify the identity of the pedestrian.

2. The method according to claim 1, characterized in that The local feature extraction module performs the following process: Among them, G v and G t is a global feature splitter with two learnable parameters, σ is the sigmoid activation function, and represents a 1×1 convolution operation, A v is the global feature segmenter G v The corresponding output, A t is the global feature segmenter G t The corresponding output is, is the global feature of the visible light mode, is the global feature of infrared mode, is the S-block local attention map of the visible light modality, is the S-block local attention map of the infrared modality, is the local feature of the visible light mode, It is the local feature of infrared mode.

3. The method according to claim 2, characterized in that The local feature alignment module performs: Grouping the local features of the visible light mode and the local features of the infrared mode to obtain s combinations; Each set of cross-modal local features is fed into different convolution modules conv i =[conv1,conv2,…,conv s ], and obtain the local features of visible light modality alignment Aligning local features with infrared modality It is expressed as: Among them, the convolution module conv consists of a convolution block with a convolution kernel size of 1 and a nonlinear activation function ReLU.

4. The method according to claim 3, characterized in that The feature collaborative learning module performs the following process: Global features of visible light modality and local features of visible light Perform global average pooling and map it to the same feature space through a linear layer to obtain and It is expressed as: Where g(·) represents the global pooling operation and Linear represents the linear layer; right and Perform the dot product operation to obtain the aggregation relationship graph E, which is expressed as: Calculate response characteristics And by and Perform spatial fusion to obtain enhanced local features It is expressed as: Among them, ω i are learnable weights.

5. The method according to claim 4, characterized in that For the local feature extraction module, a feature segmentation loss is designed Among them, S represents the number of local features, is the segmentation loss of the local features of the visible light modality, is the segmentation loss of local features of infrared modality.

6. The method according to claim 3, characterized in that For the local feature alignment module, a local group alignment loss is designed It is expressed as: Among them, i and j represent the group numbers of different cross-modal local features, and there are s groups in total.

7. The method according to claim 4, characterized in that The linear layer is a fully connected layer.

8. The method according to claim 1, characterized in that: The global feature extraction module is a pre-trained deep neural network.

9. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Near infrared-visible light cross-modal double-current pedestrian re-identification method and system

    CN114220124A

  • Cross-modal pedestrian re-identification method and system based on multi-feature learning

    CN114495010A

  • Confrontation learning cross-modal pedestrian re-identification method based on global and local features

    CN115063832A

  • Cross-modal person re-identification method and device

    WO2022027986A1

Cited By

  • Cross-modal pedestrian re-identification method and system based on multi-frequency domain features

    CN120472506A

  • Video target tracking method and device, medium and equipment

    CN120598992A

  • Power line identification method and device based on image fusion, equipment and medium

    CN120689709A

  • Weak supervision infrared visible light pedestrian re-identification method based on heterogeneous expert joint learning

    CN120953669A

  • Unmanned aerial vehicle detection method and system based on generative multi-modal fusion

    CN121280957A