Object Re-identification Method Based on Generative Multimodal Image Fusion
Through a method based on generative multimodal image fusion, a cross-modal image generation network and a dual multimodal feature fusion module are used, combined with contrast learning and attention mechanisms, the modal difference and noise problems in cross-modal target re-identification are solved, and a more accurate pedestrian target re-identification effect is achieved.
Patent Information
- Application Number
- CN202310616550.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-29
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-05-29
AI Technical Summary
The prior art is difficult to effectively eliminate modal differences in cross-modal target re-identification, resulting in the generated images lacking authenticity, loss of pedestrian identity information, and noise problems.
Using a method based on generative multimodal image fusion, paired images are generated through a cross-modal image generation network, and fusion features are used based on block-based dual multimodal feature fusion modules to generate more accurate feature representations.
Effectively eliminate cross-modal differences, retain pedestrian identity information, improve the quality of the generated image, solve the problem of modal differences and noise, and achieve more accurate pedestrian target re-identification.
Smart Images

Figure CN116824625B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target re-identification, and in particular to a target re-identification method based on generative multimodal image fusion. Background Art
[0002] Object re-identification is a computer vision task that aims to match images of the same object taken by different cameras in different occasions. Cross-Modality Person Re-identification mainly studies the matching between infrared images and visible light images: given a visible light (infrared) image of a pedestrian, it is necessary to search for the corresponding infrared (visible light) image from another modality.
[0003] In recent years, with the improvement of public safety needs, target re-identification has become one of the hot topics in the field of video surveillance and has great research significance. However, in most real-world scenarios, it is not enough to simply retrieve images obtained during the day. Thanks to the fact that most existing cameras support automatic switching between visible light and infrared modes, cross-modal target re-identification has received more and more attention. It aims to match images of people taken by visible light and infrared cameras, but due to the large differences between images of different modalities, it is difficult to achieve in practical applications.
[0004] Therefore, the existing solutions are mainly divided into two categories: non-generation-based methods and generation-based methods. Most existing models are based on non-generation models, which mainly reduce the differences between modalities through representation learning and metric learning. By extracting modal shared features, the information of different modalities is mapped to the same feature space and then discriminative features are learned. Usually, these methods use a two-stream structure to extract features. Although these methods have made great improvements, they only focus on the modal shared features and ignore the modal specific information in a single modality. They are also of great value for identifying pedestrians. The second type of method is the generation-based method. This type of method mainly performs image conversion based on the GAN (Generative Adversarial Network) model, converting visible light images into corresponding infrared images or converting infrared images into corresponding visible light images. The modal differences are effectively alleviated through pixel alignment and feature alignment. Although most methods have achieved modal conversion, they do not utilize specific information in the modality and fail to fully utilize and fuse the information of each modality. At the same time, since visible light images and near-infrared images have independent data and different distribution characteristics, there are data differences between the images converted from each other and the real images. The images generated by the generative adversarial network will inevitably contain noise, which will affect the feature extraction of subsequent networks, especially the conversion of infrared images to visible light images. Therefore, how to eliminate cross-modal differences and learn more accurate feature representations is the main problem we need to solve.
[0005] As a research problem with important application value, the object re-identification problem has attracted widespread attention in the industrial image field. In particular, based on the rapid development of deep learning technology in recent years, many model algorithms using convolutional neural networks have emerged.
[0006] The model structure of JSIA-ReID, a target re-identification method in the prior art, includes two parts: a cross-modal paired image generation module and a feature alignment module. In the cross-modal paired image generation module, images can be decomposed and represented as modality-invariant features and modality-specific features. For unpaired images, paired images can be generated by separating and exchanging their style information. The two generated images have the same content information (such as posture, etc.) and different style styles (such as clothing color, etc.). The generation module includes three encoders and two decoders. The three encoders are used for feature separation to decompose visible light and infrared images into modality-invariant and visible light (infrared) modality-specific features. After that, the decoder takes the modality-invariant features and modality-specific features as input. By decoding the features, cross-modal paired images can be generated. In the feature alignment module, an encoder is first used, and the weights are shared with the modality-invariant encoder to retain the modality-invariant features. However, performing global set-level alignment only between the entire visible light and infrared sets may lead to misalignment of some instances. To overcome this problem, instance-level alignment is performed using cross-modal paired images generated by the generation module, and the paired images of the two modalities are aligned by minimizing the KL divergence, and the modality difference is further reduced by minimizing the distance of each pair of images. By jointly training the generation module and the feature alignment module, modality alignment and identity discriminant features are learned.
[0007] The disadvantages of the target re-identification method JSIA-ReID in the above prior art include:
[0008] 1. The images generated by this method lack authenticity and will also lose the identity information of pedestrians. The images generated by the model have inconsistent colors.
[0009] 2. Although this method model generates cross-modal paired images, it does not utilize its modality-specific features and loses information.
[0010] 3. This method does not make good use of the paired images generated by the generative model. Even though a few studies have combined visible light and infrared images to bridge the cross-modal differences, they use simple fusion methods such as direct addition, which makes the noise in the generated image more obviously introduced and makes the quality of the generated image more sensitive. Summary of the invention
[0011] The embodiments of the present invention provide a target re-identification method based on generative multimodal image fusion to achieve effective pedestrian target re-identification in multimodal images.
[0012] In order to achieve the above object, the present invention adopts the following technical scheme.
[0013] A target re-identification method based on generative multimodal image fusion, comprising:
[0014] Generate an image of another modality of an input image using a cross-modal image generation network to obtain a paired image, wherein the input image includes a visible light image or an infrared image;
[0015] Constructing and training a block-based dual multimodal feature fusion module, inputting the paired images into the trained dual multimodal feature fusion module to obtain fusion features of the paired images;
[0016] Based on the fusion features of the paired images, the difference features between the paired images are measured using the cosine distance, and the difference features are compensated.
[0017] Preferably, the method of using a cross-modal image generation network to generate an image of another modality of an input image to obtain a paired image, wherein the input image includes a visible light image or an infrared image, comprises:
[0018] The cross-modal image generation network based on contrastive learning uses an encoder to separate features, uses two decoders to generate paired images, uses a visible light image or an infrared image as an input image, inputs the input image into the cross-modal image generation network, and the cross-modal image generation network uses the encoder to obtain modality-specific features and modality-shared features of the input image, and then inputs the modality-specific features and modality-shared features into two decoders respectively, and fuses the output images of the two decoders to obtain an image of the other modality of the input image;
[0019] When the input image is a visible light image, the modality-specific features and modality-shared features of the visible light image are used to generate an infrared image of another modality; when the input image is an infrared image, the modality-specific features and modality-shared features of the infrared image are used to generate a visible light image of another modality, and the input image and the image of the other modality of the input image are combined into a paired image.
[0020] Preferably, the method further comprises:
[0021] The generation process of the image of another modality of the input image is constrained by a patch-based cross-modal contrast loss, and the input and output images are matched at corresponding positions by calculating the patch-based cross-modal contrast loss;
[0022] The patch-based cross-modal contrast loss is shown in Formula 4.
[0023]
[0024] L is the number of spatial positions of the feature map, and the relevant position features of the input image corresponding to the generated image are The irrelevant position feature is A noise contrast estimation framework is used to relate input and output, l represents the cross entropy loss, and the variable yes spatial location, is considered a positive sample. These locations are considered negative samples.
[0025] Preferably, the construction and training of the block-based dual multimodal feature fusion module includes:
[0026] Input the multimodal pedestrian image of the training set into the block-based dual multimodal feature fusion module, and after some image preprocessing operations, forward propagate the data through the designed network model to obtain the classification result of the image, then calculate the loss, use the loss for back propagation, update the model weight, repeat this process until the set number of iterations is reached, and obtain the preliminarily trained dual multimodal feature fusion module;
[0027] During the testing phase, the data of the test set is loaded, and the neural network layer of the classification part is removed through the trained dual multimodal feature fusion module to directly obtain the test sample features, calculate and compare the feature similarity, complete the retrieval process, and then calculate the evaluation index to determine the model performance. If it does not meet the expected requirements, return to the training phase for further adjustment and training. If the expected performance has been achieved, save the dual multimodal feature fusion module weights to obtain the final trained dual multimodal feature fusion module.
[0028] Preferably, the step of inputting the paired images into a trained dual multimodal feature fusion module to obtain fusion features of the paired images comprises:
[0029] The paired images are input into the trained dual multimodal feature fusion module, which divides the paired images into P equal parts. Each part first passes through the first multi-branch channel attention module M, where M combines the channel attention of global and local features. The paired images are added and sent to two branches. The two branches extract global and local attention respectively. The local attention branch adopts two layers of 1×1 convolution operation. The global attention branch adds a global average pooling operation compared to the local attention branch. After adding the global and local attentions, a Sigmoid operation is performed to obtain an initial fusion weight, and the initial fusion weight is used to calculate the primary fusion feature PF;
[0030] The second attention module is used to fuse the input partial-level images, the primary fusion feature PF is input into the second attention module to obtain the partial-level fusion feature FF, and P partial-level fusion features FF are concatenated to obtain the fusion feature of the paired images.
[0031] It can be seen from the technical solutions provided by the above embodiments of the present invention that the present invention utilizes the information of the lost modality-specific features to fundamentally solve the modality difference problem. In the image generation module, in addition to focusing on the conversion of image styles, contrastive learning is used to associate the information of the corresponding positions of the input and output images, and the identity information of pedestrians is retained, so that pedestrian targets can be effectively re-identified in multimodal images.
[0032] Additional aspects and advantages of the present invention will be given in part in the following description, which will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0034] Figure 1 A processing flow chart of a method based on generative multimodal image fusion provided by an embodiment of the present invention;
[0035] Figure 2 A structural diagram of a cross-modal image generation network proposed in an embodiment of the present invention;
[0036] Figure 3 A training flow chart of a multimodal image fusion model provided by an embodiment of the present invention;
[0037] Figure 4 A structural diagram of a dual multimodal feature fusion module based on block provided by an embodiment of the present invention;
[0038] Figure 5 A schematic diagram of a test mode based on multimodal information completion provided by an embodiment of the present invention;
[0039] Figure 6 A processing flow chart of an image pedestrian target re-identification method based on generative multimodal image fusion provided by an embodiment of the present invention; DETAILED DESCRIPTION
[0040] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention.
[0041] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0042] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.
[0043] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0044] The purpose of the present invention is to fundamentally solve the modal difference problem of cross-modal target re-identification. Through the idea of modal compensation, a method based on generative multimodal image fusion is designed. The method uses a designed cross-modal image generation network based on contrastive learning to generate a corresponding image of another modality by extracting modality-specific features and modality-shared features. A patch-based cross-modal contrast loss is introduced to ensure the consistency of corresponding position information between the generated image and the input image. At the same time, a block-based dual multimodal feature fusion module is designed to fully fuse the information of the generated image and the original image, horizontally segment the image, and realize fine-grained fusion of the image. Multi-scale features are fused by combining local and global features, and an attention mechanism is used to enhance the discriminative features of the image, ultimately achieving a unified feature representation.
[0045] The processing flow of a method based on generative multimodal image fusion provided by an embodiment of the present invention is as follows: Figure 1 As shown, the processing steps include the following:
[0046] Step S10: Generate an image of another modality of the input image using a cross-modal image generation network to obtain a paired image. The input image includes a visible light image or an infrared image.
[0047] A cross-modal image generation network based on contrastive learning is used to generate paired images, and a patch-based cross-modal contrastive loss is designed to complete the paired image information of one modality for the image of another modality in the image space.
[0048] Step S20: construct and train a block-based dual multimodal feature fusion module, input the paired images into the trained dual multimodal feature fusion module, and obtain fusion features of the paired images.
[0049] The block-based dual multimodal feature fusion module is implemented using a series of 1×1 convolution operations and uses a multi-branch fusion strategy in spatial and channel dimensions to extract richer fusion feature representations of paired images.
[0050] Step S30: measuring the difference features between the paired images using the cosine distance based on the fusion features of the paired images, and compensating for the difference features.
[0051] A test method based on multimodal information completion is implemented. During the matching process, cross-modal images can be generated for query images (query) and search gallery images (gallery), and information representation can be aligned to eliminate the semantic gap of cross-modal matching.
[0052] Specifically, the above step S10 includes: the embodiment of the present invention proposes a structure of a cross-modal image generation network as follows Figure 4 The above cross-modal image generation network uses three encoders to separate features and two decoders to generate paired images.
[0053] In the cross-modal image generation network based on contrastive learning, the input image is a visible light image or an infrared image. First, the encoder is used to obtain the modality-specific features and modality-shared features of the input image. Then, the modality-specific features and modality-shared features extracted from the two modalities are input into two decoders respectively. The output images of the two decoders are fused to obtain the image of the other modality of the input image.
[0054] When the input image is a visible light image, the modality-specific features and modality-shared features of the visible light image are used to generate an infrared image of another modality; when the input image is an infrared image, the modality-specific features and modality-shared features of the infrared image are used to generate a visible light image of another modality, and the input image and the image of the other modality of the input image are combined into a paired image.
[0055] like Figure 4 As shown in the lower right, the embodiment of the present invention designs a new patch-based cross-modality contrastive loss (PCMC) for image generation.
[0056] loss) is used to constrain the image generation process. For pedestrian images, the information of the body parts in the generated image is more closely related to the body parts at the corresponding positions in the input than the background parts. For the generated image and the original image, we input them into the encoder of the previous corresponding generator to extract features, and then generate features through a two-layer MLP (Multi-LayerPerceptron) network. After that, we obtain the relevant position features and the irrelevant position features, and match the input and output images at the corresponding positions by calculating the patch-based cross-modal contrast loss.
[0057]
[0058]
[0059] Taking the generation of color images of infrared images as an example, Represents the generated color image. The input data of Formula 1 is the original image and the corresponding cross-modal generated image. For extracting modality-specific features of visible images ( For extracting modality-specific features of infrared images). F is a two-layer MLP network, and we use the MLP network (F) to extract features. Each spatial position in the new feature map obtained by F is a vector with a dimension equal to the number of channels.
[0060] The patch-based cross-modal contrast loss is shown in Formula 4, where L is the number of spatial positions of the feature map, and the relevant position features of the input image corresponding to the generated image are The irrelevant position feature is We use the noise contrast estimation framework to associate input and output, l represents the cross entropy loss, and the cross entropy loss (Formula 3) is used to calculate the association between input and output data. yes spatial location. is considered a positive sample, and These positions are considered negative samples. Subsequently, the PCMC loss is calculated based on these samples. The input of Formula 4 is these sample features, and the sampling points are merged for calculation. In VI-ReID, pedestrian image translation needs to ensure identity consistency, and contrastive learning can help maintain content consistency and reduce noise. Therefore, we designed the PCMC loss to impose specific constraints on the generation process. The significance of this module is to make the generated image have the original identity information without introducing noise information, which can solve the problem of color inconsistency.
[0061] The output image of the cross-modal image generation network is the image of the other modality of the input image, that is, pairs of images are obtained, the real visible light image and the generated pseudo infrared image, the real infrared image and the generated pseudo visible light image.
[0062] Specifically, the above step S20 includes: a training process of a generative dual multimodal feature fusion module provided by an embodiment of the present invention is as follows: Figure 3 As shown, the processing includes the following:
[0063] First is the model training phase, where the multimodal pedestrian images of the training set are input, and some image preprocessing operations are performed, including data standardization, image size change, random horizontal flipping, random cropping and random erasing, etc. Then the data is forward propagated through the designed network model to obtain the image classification result, and then the loss is calculated, and the loss is used for back propagation to update the model weights. This process is repeated until the set number of iterations is reached. In the testing phase, the data of the test set is loaded, and the neural network layer of the classification part is removed through the trained model to directly obtain the test sample features, calculate and compare the feature similarity, complete the retrieval process, and then calculate the evaluation index to determine the model performance. If it does not meet the expected requirements, it will return to the training phase for further adjustment and training. If the expected performance has been achieved, the model weights will be saved, and the entire technical invention process will be completed to obtain the final solution.
[0064] When the multimodal feature fusion module fuses two paired images, unlike the previous simple channel splicing or direct addition method, the fusion strategy of the present invention can overcome the semantic inconsistency problem by introducing channel attention and help learn more discriminative features. Figure 4 A block-based dual multimodal feature fusion module is provided in an embodiment of the present invention, such as Figure 4As shown in the figure, the paired images obtained by the above cross-modal image generation network are first divided into P parts. Each part first passes through the first multi-branch channel attention module M to improve the input quality of the second attention module, where M combines the channel attention of global and local features. The paired images are added and sent to two branches to extract global and local attention respectively. The local attention branch adopts two layers of 1×1 convolution operation. The global attention branch adds a global average pooling operation compared to the local attention branch. The global and local attentions are added and then Sigmoid operation is performed to obtain an initial fusion weight. The initial fusion weight can be used to calculate the primary fusion feature PF. After that, another identical attention module is used to fuse the input partial-level image, that is, the second attention module. The above primary fusion feature PF is input into the second attention module, and finally the partial-level fusion feature FF is obtained. The fusion features FF of the P parts are spliced together to obtain the fusion features of the above paired images.
[0065] Specifically, the above step S30 includes: Figure 5 A schematic diagram of a test mode based on multimodal information completion provided by an embodiment of the present invention implements a test method based on multimodal information compensation without changing the test mode of the data set. Figure 5 As shown in the lower side (b), compared with the traditional cross-modal target re-identification test method ( Figure 5 (a)) The difference is that we let all images in the query set and search library images pass through the cross-modal image generation network and the block-based dual multimodal feature fusion module to achieve data compensation and unification, and use the cosine distance to measure the difference features between the paired images based on the fusion features of the paired images, and compensate for the difference features, thereby solving the modality difference problem.
[0066] Based on the above generative multimodal image fusion model, a processing flow chart of an image pedestrian target re-identification method based on generative multimodal image fusion provided by an embodiment of the present invention is as follows: Figure 6 As shown in the figure, it includes the following stages: image preprocessing, cross-modal image generation, multimodal feature fusion, feature extraction, model iteration optimization, and model testing and evaluation. The specific operations of each stage are explained in detail below.
[0067] Input to the algorithm:
[0068] 1. Cross-modal image data: including training images (train), query images (query) and search gallery (gallery). The training images are used to train the above-mentioned generative multimodal image fusion model to extract features, and the search images and search gallery are used to verify the performance of the above-mentioned generative multimodal image fusion model.
[0069] 2. Model algorithm hyperparameters: including image cropping size, batch size in training, number of iterations and learning rate, number of image blocks P, balance coefficient λ of patch-based cross-modal contrast loss, etc.
[0070] Output of the algorithm:
[0071] The parameter weights of the trained generative multimodal image fusion model that meets the performance evaluation criteria are obtained. Model testing process: The input of the model is the query image. Through the generative multimodal image fusion model, we can obtain image pairs and perform feature fusion on the obtained paired images to achieve data compensation and unification, and output the final feature embedding.
[0072] Execution steps:
[0073] 1. Image preprocessing stage
[0074] Step 1-1: Load the image data of the cross-modal dataset (including the training set and the test set) into the GPU memory; the image data in the training set and the test set include visible light images and infrared images.
[0075] Step 1-2: Standardize the images of the training set and the test set, scale the pixel value range of the image to between 0 and 1, crop the image according to the set size, and use data enhancement operations such as random horizontal flipping and random cropping as appropriate;
[0076] Step 1-3: Organize the image data into batches according to the set batch size for input into the subsequent model algorithm.
[0077] 2. Cross-modal paired image generation stage
[0078] Step 2-1: For the preprocessed visible light images and infrared images in the training set, use different encoders to extract their modality-specific features and modality-shared features;
[0079] Step 2-2: The obtained modality-specific features and modality-shared features are fed into different decoders to generate corresponding visible light images and infrared images;
[0080] Step 2-3: Calculate adversarial loss, reconstruction loss, cycle consistency loss, and cross-modal contrast loss (Formula 4), perform weighted fusion on these losses, and obtain the loss value;
[0081] Step 2-4: Using the PyTorch deep learning framework, backpropagation can be performed starting from the final calculated loss value, the gradient values of the parameters in the model can be automatically calculated, and the optimizer (such as Pytorch's Adam optimizer) can be used to update the learnable parameter values of the model algorithm. All the above execution steps are repeated before the model reaches the number of rounds set by the hyperparameters. After the number of training rounds is reached, the model training process is stopped to finally obtain the generated cross-modal paired images.
[0082] 3. Multimodal feature fusion stage
[0083] Step 3-1: Divide the obtained paired images into P partial images respectively;
[0084] Step 3-2: Input each part of the divided image into the first multi-branch channel attention module to obtain preliminary fusion features;
[0085] Step 3-3: Input the obtained preliminary fusion features into the second identical attention module to obtain the final part-level fusion features;
[0086] Step 3-4: Concatenate the fusion features of different parts to obtain the fusion features of the final image.
[0087] 4. Feature extraction stage
[0088] Step 4-1: Input the fused features into the backbone network of cross-modal person re-identification for feature extraction;
[0089] Step 4-2: Calculate the classification loss (ID loss);
[0090] Step 4-3: Calculate triplet loss;
[0091] Step 4-4: Add up the losses to get the value of the total loss.
[0092] 5. Model Iteration Optimization Phase
[0093] Step 5-1: The code implementation is based on the PyTorch deep learning framework, which can perform backpropagation based on the final calculated loss value and automatically calculate the gradient value of the parameters in the model;
[0094] Step 5-2: Using the gradient calculated in the previous step, use an optimizer (such as Pytorch's Adam optimizer) to update the learnable parameter values of the model algorithm;
[0095] Step 5-3: Repeat all the above execution steps until the model reaches the number of rounds set by the hyperparameters, and stop the model training process after reaching the number of training rounds.
[0096] 6. Test and Evaluation Phase
[0097] Step 6-1: Read the pedestrian images of the test set, load them into the GPU memory, and perform the same standardization operations as the training phase (note that data enhancement operations such as random horizontal flipping are not required during testing);
[0098] Step 6-2: Use the Cumulative Matching Characteristics (CMC) and Mean Average Precsion (mAP) evaluation indicators commonly used in pedestrian re-identification to preliminarily evaluate the quality of the model by evaluating the calculated indicator values.
[0099] Step 6-3: If the evaluation result does not meet the requirements, you need to adjust the model's hyperparameters, return to the first step of the execution step, and re-train the model. If the evaluation result meets the requirements, you can save the model weights and obtain the solution to this visible light infrared cross-modal pedestrian re-identification task.
[0100] Before using the current technical invention, technical personnel are required to configure the relevant links, including installing the Linux operating system, Python 3.7 (and above) development environment, and PyTorch1.7 (and above) deep framework. Because the algorithm used in the present invention is a model algorithm based on deep learning, it is recommended to perform the model training process in a GPU (graphics processing unit) environment, and it is necessary to install the GPU version of Pytorch1.7 (and above) and the corresponding version of the CUDA parallel computing architecture.
[0101] It can be seen from the description of the above implementation methods that the execution subject of the training method of the target re-identification model of this embodiment can be a training device for the target re-identification model, which can be implemented by software and / or hardware, and the device can be configured in an electronic device, and the electronic device may include but is not limited to a terminal, a server, etc. The multiple images in this embodiment may have a visible light modality and an infrared modality. In practical applications, an image acquisition device (for example, a camera) can be used to collect visible light images or video frames during the day and infrared images or video frames at night, so that multiple images with multiple modalities can be obtained. In this embodiment, by obtaining multiple images having corresponding multiple modalities and corresponding multiple labeled target categories, the re-identification model trained can fully exploit the features in the images of multiple modalities, and can enhance the accuracy of image matching under different modalities, thereby improving the effect of cross-modal target re-identification. The technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, a disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0102] In summary, compared with the existing cross-modal target re-identification method for extracting modality shared features, the method of the embodiment of the present invention utilizes the information of the lost modality specific features and fundamentally solves the modality difference problem.
[0103] In the image generation module, in addition to focusing on the conversion of image styles, the method of the embodiment of the present invention uses contrastive learning to associate the information of corresponding positions of the input and output images, thereby retaining the identity information of the pedestrians.
[0104] In the multimodal feature fusion stage, the method of the embodiment of the present invention realizes fine-grained fusion based on the idea of block segmentation, and introduces channel attention to overcome the problem of semantic inconsistency and learn more discriminative feature representation.
[0105] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.
[0106] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.
[0107] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0108] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A target re-identification method based on generative multimodal image fusion, It is characterized in that include: Generate an image of another modality of an input image using a cross-modal image generation network based on contrastive learning to obtain a paired image, wherein the input image includes a visible light image or an infrared image; Constructing and training a block-based dual multimodal feature fusion module, inputting the paired images into the trained dual multimodal feature fusion module to obtain fusion features of the paired images; Measuring the difference features between the paired images using the cosine distance based on the fusion features of the paired images, and compensating for the difference features; The cross-modal image generation network based on contrastive learning uses an encoder to separate features, uses two decoders to generate paired images, uses a visible light image or an infrared image as an input image, inputs the input image into the cross-modal image generation network, and the cross-modal image generation network uses the encoder to obtain modality-specific features and modality-shared features of the input image, and then inputs the modality-specific features and modality-shared features into two decoders respectively, and fuses the output images of the two decoders to obtain an image of the other modality of the input image; Inputting the paired images into the trained dual multimodal feature fusion module to obtain fusion features of the paired images includes: The paired images are input into the trained dual multimodal feature fusion module, which divides the paired images into P equal parts. Each part first passes through the first multi-branch channel attention module M, where M combines the channel attention of global and local features. The paired images are added and sent to two branches. The two branches extract global and local attention respectively. The local attention branch adopts two layers of 1×1 convolution operation. The global attention branch adds a global average pooling operation compared to the local attention branch. After adding the global and local attentions, a Sigmoid operation is performed to obtain an initial fusion weight, and the initial fusion weight is used to calculate the primary fusion feature PF; The second attention module is used to fuse the input partial-level images, the primary fusion feature PF is input into the second attention module to obtain the partial-level fusion feature FF, and P partial-level fusion features FF are concatenated to obtain the fusion feature of the paired images.
2. The method according to claim 1, It is characterized in that Generate an image of another modality of an input image using a cross-modal image generation network to obtain a paired image, wherein the input image includes a visible light image or an infrared image, including: When the input image is a visible light image, the modality-specific features and modality-shared features of the visible light image are used to generate an infrared image of another modality; when the input image is an infrared image, the modality-specific features and modality-shared features of the infrared image are used to generate a visible light image of another modality, and the input image and the image of the other modality of the input image are combined into a paired image.
3. The method according to claim 2, It is characterized in that The method further comprises: The generation process of the image of another modality of the input image is constrained by a patch-based cross-modal contrast loss, and the input and output images are matched at corresponding positions by calculating the patch-based cross-modal contrast loss; The patch-based cross-modal contrast loss is shown in Formula 4. L is the number of spatial positions of the feature map, and the relevant position features of the input image corresponding to the generated image are The irrelevant position feature is A noise contrast estimation framework is used to relate input and output, l represents the cross entropy loss, and the variable yes spatial location, is considered a positive sample. These locations are considered negative samples.
4. The method according to claim 3, It is characterized in that The construction and training of the block-based dual multimodal feature fusion module includes: Input the multimodal pedestrian image of the training set into the block-based dual multimodal feature fusion module, and after some image preprocessing operations, forward propagate the data through the designed network model to obtain the classification result of the image, then calculate the loss, use the loss for back propagation, update the model weight, repeat this process until the set number of iterations is reached, and obtain the preliminarily trained dual multimodal feature fusion module; During the testing phase, the data of the test set is loaded, and the neural network layer of the classification part is removed through the trained dual multimodal feature fusion module to directly obtain the test sample features, calculate and compare the feature similarity, complete the retrieval process, and then calculate the evaluation index to determine the model performance. If it does not meet the expected requirements, return to the training phase for further adjustment and training. If the expected performance has been achieved, save the dual multimodal feature fusion module weights to obtain the final trained dual multimodal feature fusion module.
Citation Information
Patent Citations
Cross-modal pedestrian re-identification method based on multi-modal image style conversion
CN111539255A
Near infrared-visible light cross-modal double-current pedestrian re-identification method and system
CN114220124A