A visible light and infrared pedestrian re-identification method based on feature decoupling

By introducing a feature decoupling layer in the deep learning model, decoupling identity features and redundant features, and adopting an end-to-end training method, the problem of high difficulty in model training and low recognition rate in the existing technology is solved, and a more efficient pedestrian re-identification effect is achieved.

CN117496552BActive Publication Date: 2025-05-09SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311342566.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-17
Publication Date
2025-05-09
Estimated Expiration
2043-10-17

AI Technical Summary

Technical Problem

The prior art has the problem that the image generation process depends on the quality of the generated image in visible light and infrared pedestrian re-recognition, which leads to high difficulty in model training, low recognition rate, and high resource consumption in multi-stage training.

Method used

A deep learning model based on feature decoupling is adopted, and identity features and redundant features are decoupled through the feature extraction layer, feature decoupling layer and classifier layer. An end-to-end training method is adopted, without involving the image generation process, and the training stage is reduced.

Benefits of technology

The model's optimization effect on pedestrian characteristics is improved, the compactness and separability of identity characteristics is enhanced, redundant information interference is reduced, and the accuracy and computing efficiency of pedestrian re-identification are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117496552B_ABST
    Figure CN117496552B_ABST
Patent Text Reader

Abstract

The present invention discloses a visible light and infrared pedestrian re-identification method based on feature decoupling. The method comprises: collecting a target area image, wherein the target area image comprises a visible light image modality and an infrared image modality; inputting the target area image into a trained deep learning model to obtain a pedestrian recognition result; wherein the deep learning model comprises a feature extraction layer, a feature decoupling layer and a classifier layer, wherein the feature extraction layer is used to extract pedestrian features from the target area image, the feature decoupling layer is used to decouple identity features and redundant features from the extracted pedestrian features, and the classifier layer is used to identify the target pedestrian based on the decoupled identity features. The present invention can accurately decouple redundant features and identity features from pedestrian features, thereby enhancing the optimization effect of features and improving the accuracy of pedestrian recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a visible light and infrared pedestrian re-identification method based on feature decoupling. Background Art

[0002] Visible light and infrared pedestrian re-identification aims to match pedestrian images of two modalities and can be applied to nighttime intelligent monitoring systems. With the development of monitoring technology, cameras usually have visible light mode and infrared mode, and the two modes can be automatically switched during the day and night. The spectral imaging principles of these two modes are different, and the images collected also have a heterogeneous relationship. Specifically, visible light images contain color information of three channels: red, green, and blue, while infrared images have only one channel containing information about the intensity of infrared light. Therefore, visible light and infrared pedestrian re-identification not only has to deal with the feature changes of the same pedestrian in different images under a single modality, such as different background interference, different pedestrian postures, changes in viewing angles, changes in lighting, etc., but also needs to deal with cross-modal differences caused by the natural difference between the reflectivity of the visible spectrum and the emissivity of the infrared spectrum.

[0003] In order to address the above problems and reduce the interference of redundant information, some existing solutions use different encoders for visible light and infrared pedestrian images, decouple identity information and modal redundant information, and then use decoders to unify cross-modal image representations. However, these methods all involve the process of image generation and rely heavily on the quality of the generated images. The generated images require high-dimensional features to be reconstructed pixel by pixel. These strict generation conditions increase the difficulty of model training and make it difficult for the model to reconstruct pedestrian images well. Therefore, in the process of image generation, some important discriminant information will be lost and additional noise will be introduced. For example, information such as the outline of pedestrians and the texture of clothes will change significantly or be lost.

[0004] In addition, there are some existing solutions that first extract identity-related features with high redundancy, and then separate modality-specific redundant information, such as color information of visible light images and thermal information of infrared images. However, these methods do not fully consider the interference caused by background information, and they all require multiple stages of training, which consumes more training time and larger space resources under the same conditions, and the algorithm complexity is high.

[0005] In summary, existing methods usually involve a complex image generation process and a multi-stage training process, which reduces computational efficiency and model training efficiency. Summary of the invention

[0006] The purpose of the present invention is to overcome the defects of the prior art and provide a visible light and infrared pedestrian re-identification method based on feature decoupling. The method comprises the following steps:

[0007] Acquiring a target area image, wherein the target area image includes a visible light image modality and an infrared image modality;

[0008] Inputting the target area image into a trained deep learning model to obtain a pedestrian recognition result;

[0009] Among them, the deep learning model includes a feature extraction layer, a feature decoupling layer and a classifier layer. The feature extraction layer is used to extract pedestrian features from the target area image, the feature decoupling layer is used to decouple identity features and redundant features from the extracted pedestrian features, and the classifier layer is used to identify the target pedestrian based on the decoupled identity features.

[0010] Compared with the prior art, the advantage of the present invention is that, in order to solve the problem that the network model is disturbed by redundant information such as lighting changes and background changes when extracting pedestrian features, resulting in a low model recognition rate, a visible light and infrared pedestrian re-identification method based on feature decoupling is provided. Identity features and redundant features are decoupled by designing a feature decoupling module. This method of separating redundant features enhances the compactness of the same category and the separability of different categories in the identity features. The feature decoupling module does not adopt an encoding-decoding structure, does not involve the image generation process, and adopts an end-to-end training method, which does not require setting multiple training stages, thereby improving the model's optimization effect on pedestrian features.

[0011] Further features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the present invention with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0013] Figure 1 is a flow chart of a visible light and infrared pedestrian re-identification method based on feature decoupling according to an embodiment of the present invention;

[0014] Figure 2 is an architectural diagram of a deep learning model according to an embodiment of the present invention;

[0015] Figure 3 is an example of a heat map visualization according to an embodiment of the present invention;

[0016] Figure 4 is an example of feature distribution visualization according to an embodiment of the present invention. DETAILED DESCRIPTION

[0017] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.

[0018] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.

[0019] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered as part of the specification.

[0020] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0021] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0022] See also Figure 1 As shown, the provided visible light and infrared pedestrian re-identification method based on feature decoupling includes the following steps:

[0023] Step S110, constructing a deep learning model, which generally includes a feature extraction layer, a feature decoupling layer and a classifier layer.

[0024] Combination Figure 2 As shown, the constructed deep learning model includes a feature extraction layer 10, a feature decoupling layer 20 and a classifier layer 30 in sequence. The feature extraction layer 10 is used to extract pedestrian features in the surveillance image. The feature decoupling layer 20 is used to decouple identity features and redundant features from pedestrian features. The classifier layer 30 is used to accurately classify the categories of identity features.

[0025] For example, an existing deep neural network is selected as the feature extraction layer, and the deep neural network type can be a convolutional neural network or a recursive neural network and other types.

[0026] In one embodiment, the feature decoupling layer includes two modules: instance normalization IN and channel mask M. For example, instance normalization IN can be expressed as:

[0027]

[0028] Among them, F represents the pedestrian features obtained by the feature extraction layer, represents the normalized features after the IN operation, μ and σ represent the mean and standard deviation calculated in the spatial dimension for each independent channel of each instance feature, respectively. They are respectively learnable scaling parameters and translation parameters. From the calculation method of IN, we can see that IN can maintain the independence between the features of each instance and reduce the degree of mutual influence between different modes.

[0029] Although the IN operation can play a good role in filtering modality-specific information, it does not consider the identity correlation between different channels in each instance feature, making it difficult to enhance the overall identity discriminability in the instance feature. The residual feature R between is taken as the decoupled object, and the residual feature R is defined as:

[0030]

[0031] The residual feature R represents the original feature F and the instance normalized feature There are two main reasons for using R as the decoupling object: First, the residual feature R contains The lost identity information, secondly, if the valid identity information contained in the residual feature R and The new decoupled features are reorganized, which can avoid the unstable image generation process compared to the encoding-decoding decoupling process and solve the problems existing in the IN operation.

[0032] Considering that the residual feature R contains both the lost identity information and the removed redundant information, in order to decouple the identity information and redundant information from the residual feature, the channel mask M is used to decouple the residual feature R and adaptively separate the identity information and redundant information in R. This process is expressed as:

[0033]

[0034] R + =M×R (4)

[0035] R - =(1-M)×R (5)

[0036] Where g(R) represents the global average pooling of R. and represents the learnable parameters in two fully connected layers without setting bias, c represents the number of feature channels, δ and σ represent the ReLU activation function and the Sigmoid activation function respectively. In order to improve adaptability and reduce the number of parameters, for example, the dimension reduction ratio r is set to 16. In addition, R + It is called the positive residual feature, which represents the identity information contained in R. - It is called negative residual feature, which represents the redundant information contained in R.

[0037] Then, R + and R - Respectively Recombination to obtain identity characteristics and redundant features

[0038]

[0039]

[0040] The classifier layer uses the decoupled identity features to determine whether a specific target pedestrian exists in the image or video sequence, that is, to obtain the pedestrian re-identification result. The classifier layer can be implemented using a fully connected layer or other network architectures. In this article, the fully connected layer is used as an example for description.

[0041] Step S120, performing end-to-end training on the deep learning model using the set loss function.

[0042] During the training phase, the entire model can be trained using the cross entropy loss function and the triplet loss function. In addition, in order to better promote the decoupling effect, the feature decoupling layer is additionally trained using the adversarial decoupling loss function.

[0043] For example, the cross entropy loss function is expressed as:

[0044]

[0045] Among them, N represents the number of pedestrian images under the same modality in each training batch, m represents the number of pedestrian categories in the cross-modal training dataset, subscript i represents the index of the current batch of pedestrian images, subscript j represents the category of the pedestrian, v represents the parameters of the fully connected layer used for classification, and f represents the pedestrian features extracted from the pedestrian image through the model.

[0046] For example, the triplet loss mainly takes a triplet (anchor sample a, positive sample p and negative sample n) as input, where p and a have the same label, while n and a have different labels. After the model extracts features from the input sample, it obtains the positive sample feature pair {f a ,f p} and negative sample feature pairs {fa ,f n}, the triplet loss is to continuously drive the distance between the positive sample feature pairs to be less than a set value than the distance between the negative sample feature pairs, so as to achieve feature clustering. The calculation expression is:

[0047]

[0048] Among them, y represents the category corresponding to the feature of the input sample, ρ 1 is a predefined interval value.

[0049] Furthermore, adversarial disentanglement loss is designed to promote the disentanglement of identity features and redundant features, for example, defining the center μ of the feature distribution belonging to category c c for:

[0050]

[0051] in, represents the i-th sample feature belonging to category c, and K represents the number of pedestrian features belonging to category c under the same modality in each batch.

[0052] Then, in each training batch, the Euclidean distance between all sample features belonging to the same category and the center of their feature distribution is called the intra-class distance D intra The Euclidean distance between the feature distribution centers belonging to different categories is called the inter-class distance D inter , the calculation expressions of these two distances are:

[0053]

[0054]

[0055] Among them, P represents the number of pedestrian identities in the current training batch, and j and k represent two different identities.

[0056] According to the above definition, With R ′ After reorganization Relative to The intra-class distance D intra will be smaller, and The inter-class distance D inter Will be bigger. intra and D inter As two opposing factors, Adversarial disentanglement loss for:

[0057]

[0058] on the contrary, With R- After reorganization Relative to The intra-class distance D intra It will be bigger. The inter-class distance D inter will be smaller because It contains redundant information such as contrast, background, resolution, etc. that is irrelevant to identity, so It has weaker identity discrimination and worse clustering effect. Adversarial disentanglement loss It is expressed as:

[0059]

[0060] in addition, and The ln(1+exp(·)) used in is a monotonically increasing function, and there will be no negative loss value, which can stabilize the feature decoupling training process and reduce the difficulty of optimization. The final calculation expression of the adversarial decoupling loss is:

[0061]

[0062] The feature decoupling layer uses the adversarial decoupling loss to exert the adaptive function of the channel mask and guide the channel mask M to separate the identity information and redundant information in the residual feature R, and then separates the feature R + , R - and Recombination is performed and recombination characteristics and Different adversarial relationships realize a complete, effective and stable adversarial decoupling process. In general, the feature decoupling layer can not only decouple more robust identity discriminant features and reduce the interference of identity-irrelevant information during training, but also alleviate the modal difference between the visible light modality and the infrared modality.

[0063] In the end-to-end training process of a deep learning model, the overall loss function can be a combination of the cross entropy loss, the triplet loss, and the adversarial disentanglement loss, such as a weighted sum of the three.

[0064] Step S130: using the trained deep learning model to identify the identity of pedestrians in the acquired visible light image and infrared image.

[0065] After the model training is completed, it can be used for actual person re-identification. For example, the feature extraction layer and feature decoupling layer are retained as the trained network model, and similarity calculation is performed based on the decoupled identity features to obtain the result of person re-identification.

[0066] It should be understood that the training process of the deep learning model involved in the present invention can be carried out offline on a server or in the cloud, and real-time pedestrian identification can be achieved by embedding the trained model into an electronic device. The electronic device can be a terminal device or a server, and the terminal device includes any terminal device such as a monitoring device, a mobile phone, a tablet computer, a personal digital assistant (PDA), a car computer, a smart wearable device (smart watch, virtual reality glasses, virtual reality helmet, etc.). The server includes but is not limited to an application server or a Web server, and can be an independent server or a cluster server or a cloud server, etc.

[0067] In order to further verify the effect of the present invention, an ablation experiment was carried out and the experimental results were analyzed by visualization method.

[0068] (1) Ablation experiment

[0069] Experiments were conducted on two mainstream visible light and infrared cross-modal pedestrian re-identification datasets, which are introduced as follows.

[0070] SYSU-MM01: This dataset has a total of 491 pedestrian identities, 287,628 visible light images, and 15,792 infrared images. The training set has 395 pedestrian identities and the test set has 96 pedestrian identities. The entire dataset comes from 4 visible light cameras and 2 near-infrared cameras.

[0071] RegDB: This dataset has 412 pedestrian identities in total, and the number of visible light images and infrared images is 4120 each. There are 10 visible light images and 10 infrared images for pedestrians of the same identity. The training set and test set contain images of pedestrians of 206 identities respectively. The entire dataset comes from 1 visible light camera and 1 thermal infrared camera.

[0072] The ablation experiment uses three evaluation criteria, Rank-1, Rank-10, and mean average precision (mAP), to measure the performance of the model. All three indicators are numbers between 0 and 1. The larger the value, the higher the accuracy of person re-identification. Tables 1 and 2 show the experimental results obtained on the SYSU-MM01 and RegDB datasets by applying the method of the present invention to the baseline network with ResNet50 as the backbone. The baseline network only uses the cross entropy loss. and triplet loss Guided training does not include the feature decoupling layer and adversarial decoupling loss proposed in this invention.

[0073] Table 1 Ablation experiment results on SYSU-MM01 dataset (%)

[0074] method Rank-1 Rank-10 mAP Baseline Network 64.94 94.30 61.58 Baseline network + invention 70.97 96.43 67.46

[0075] Table 2 Ablation experiment results on RegDB dataset (%)

[0076] method Rank-1 Rank-10 mAP Baseline Network 84.96 95.53 76.65 Baseline network + invention 90.19 97.67 81.05

[0077] It can be seen from Table 1 and Table 2 that on the same training set, the present invention can significantly improve the values ​​of the three indicators of Rank-1, Rank-10 and mean average accuracy mAP of the model in the test set.

[0078] (2) Heat map visualization

[0079] The heat map visualization experiment uses the Grad-CAM method, which can calculate the activation heat map of a specific category based on the feature map output by the network model, and then use the size of the activation area in the heat map to judge the similarity between the feature map and the corresponding pedestrian image. In short, the heat map can intuitively show the model's focus area for a given pedestrian image, and these focus areas are the important basis for the model to accurately classify a given pedestrian image. This experiment selected two pedestrians. Figure 3 Show their instance normalization features in different modalities Redundant Features and identity characteristics Activation heat map of Figure 3 (a) corresponds to pedestrian 1, Figure 3 (b) Corresponding to pedestrian 2, the first row is the visible light image, and the second row is the infrared image. Pedestrian 1 is in an indoor environment, while pedestrian 2 is in an outdoor environment.

[0080] from Figure 3 It can be seen that the instance normalization feature The corresponding heat map has fewer human activation areas, which indicates Although it can alleviate the modal difference, it also loses a lot of identity information. With negative residual feature R - Recombination to obtain redundant features After that, the heat map corresponding to this feature mainly has more activation values ​​in the background area. Most of the activation areas of the corresponding heat maps are concentrated in the human body area, indicating that With positive residual feature R + After reorganization, more identity discrimination information can be obtained. Therefore, the heat map visualization experiment fully demonstrates that the feature decoupling method of the present invention can guide the model to achieve a better decoupling effect on the identity features and redundant features of the input cross-modal pedestrian images, and reduce the interference of redundant information on the identity discrimination of pedestrian features.

[0081] (3) Feature distribution visualization

[0082] The feature distribution visualization experiment uses the t-SNE method, which can reduce the dimension of high-dimensional feature vectors and convert the distance relationship between multiple high-dimensional features into a two-dimensional plane, so as to facilitate the observation and analysis of the clustering of high-dimensional features extracted by the deep learning model. This experiment randomly selects 10 categories of pedestrian images, and each category of pedestrian images has an equal number of visible light images and infrared images. After these images are input into the feature extraction network, the corresponding feature description vectors will be obtained. In the experiment, the normalized features extracted by the present invention are Redundant Features Identity The features extracted by the baseline model (containing only the feature extraction layer and no feature decoupling layer) are visualized and analyzed, corresponding to Figure 4 (a)- Figure 4 (d), where different colors represent pedestrian features of different identities, circles represent visible light features, and triangles represent infrared features. Figure 4 It can be found that although the instance normalization operation can reduce the distance between the two modes by reducing the difference between instance features, it Figure 4 The features in (a) have lost a lot of identity information. Figure 4 The redundant features in (b) contain identity-irrelevant information, so Figure 4 (a) and Figure 4 (b) does not show accurate clustering effect. Figure 4 The identity features in (c) can achieve accurate clustering using effective identity information. Compared with the features extracted by the baseline model, the feature decoupling method of the present invention can more effectively enhance the compactness of pedestrian identity features within the same category and the separability between different categories, and further reduce the modal differences between cross-modal pedestrian images in the same category.

[0083] In summary, the deep learning model of the present invention includes a feature decoupling layer, and a corresponding adversarial decoupling loss is designed, which can accurately decouple redundant features and identity features from pedestrian features, reduce the interference of redundant information on the model's extraction of pedestrian features, and can better promote the separation of identity features and redundant features, thereby enhancing the optimization effect of features. In addition, the model adopts an end-to-end training method, does not involve the image generation process, and does not require staged training, thereby improving the training efficiency of the model, ensuring the stability of the training process, and also achieving better decoupling experimental results.

[0084] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0085] Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, for example, a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (for example, a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire.

[0086] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0087] The computer program instructions for performing the operation of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, Python, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present invention.

[0088] Various aspects of the present invention are described herein with reference to the flow charts and / or block diagrams of the methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each box of the flow chart and / or block diagram and the combination of each box in the flow chart and / or block diagram can be implemented by computer-readable program instructions.

[0089] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0090] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0091] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a part of a module, a program segment or an instruction, and a part of the module, a program segment or an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.

[0092] Embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A visible light and infrared pedestrian re-identification method based on feature decoupling, comprising the following steps: Acquiring a target area image, wherein the target area image includes a visible light image modality and an infrared image modality; Inputting the target area image into a trained deep learning model to obtain a pedestrian recognition result; The deep learning model includes a feature extraction layer, a feature decoupling layer and a classifier layer, wherein the feature extraction layer is used to extract pedestrian features from the target area image, the feature decoupling layer is used to decouple identity features and redundant features from the extracted pedestrian features, and the classifier layer is used to identify the target pedestrian based on the decoupled identity features; The feature decoupling layer includes an instance normalization module IN and a channel mask module M, and decouples identity features and redundant features according to the following process: The instance normalization module IN obtains the instance normalization feature according to the following formula Where F represents pedestrian features, μ and σ represent the mean and standard deviation of each independent channel of each instance feature calculated in the spatial dimension, τ and ∈ are learnable scaling and translation parameters respectively; The channel mask module M is used to decouple the residual feature R and separate the identity information and redundant information in R, which can be expressed as: R + =M×R R - =(1-M)×R in, g(R) represents global average pooling of R. and represents the learnable parameters, δ represents the ReLU activation function, σ represents the Sigmoid activation function, R + is a positive residual feature, which represents the identity information contained in R. - It is a negative residual feature, which represents the redundant information contained in R; R + and R - Respectively After reorganization, we get the identity feature and redundant feature, which are expressed as: in, It is an identity feature. It is a redundant feature.

2. The method according to claim 1, characterized in that The overall loss function for training the deep learning model includes a cross entropy loss function and a triplet loss function, and the cross entropy loss function is expressed as: Where N represents the number of pedestrian images of the same modality in each training batch, m represents the number of pedestrian categories in the cross-modal training dataset, i represents the index of the pedestrian image in the current batch, and j represents the category of the pedestrian. represents the parameters of the fully connected layer used for classification, and f represents the pedestrian features extracted from the pedestrian image by the model; The triplet loss function is expressed as: Among them, y represents the category corresponding to the feature of the input sample, ρ1 is the predefined interval value, a is the anchor sample, p is the positive sample, n is the negative sample, {f a ,f p } is the extracted positive sample feature pair, {f a ,f n } is the extracted negative sample feature pair.

3. The method according to claim 2, characterized in that The overall loss function also contains adversarial decoupling loss, expressed as: in, About The adversarial disentanglement loss is expressed as: In each training batch, the distance between all sample features belonging to the same category and their feature distribution center is the intra-class distance D intra , the distance between the feature distribution centers belonging to different categories is the inter-class distance D inter ; in, It is expressed as: in, About Adversarial disentanglement loss.

4. The method according to claim 1, characterized in that: The feature extraction layer is a deep neural network.

5. The method according to claim 1, characterized in that The classifier layer is a fully connected layer.

6. The method according to claim 3, characterized in that The intra-class distance D intra and the inter-class distance D inter is the Euclidean distance, which is generally expressed as: Where P represents the number of pedestrian identities in the current training batch, j and k represent two different identities, and μ c is the center of the feature distribution of category c.

7. The method according to claim 6, characterized in that The center μ of the feature distribution of category c c The general expression is: in, represents the i-th sample feature belonging to category c, and K represents the number of pedestrian features belonging to category c under the same modality in each batch.

8. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

9. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.