A visible light-infrared pedestrian re-identification method and system

Through the deep network model and graph convolution network combined with attention mechanism, the cross-modal pedestrian image features are extracted and learned, and the problems of difficulty in utilization of feature information and accuracy limitation caused by the differences in cross-modal image in the prior art are solved, and higher recognition accuracy is achieved.

CN113887353BActive Publication Date: 2025-05-06SHANDONG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111107235.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-05-06
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

The existing visible-infrared pedestrian re-identification technology faces the problems of difficulty in using feature information and accuracy limitations caused by cross-modal image differences.

Method used

The deep network model is used to combine the graph convolution network and attention mechanism, and the convolution layer is connected through branch structures and the weight is dynamically adjusted, image features are extracted and learned to achieve cross-modal feature alignment and matching.

Benefits of technology

The discrimination ability of cross-modal pedestrian images is improved, and the recognition accuracy of infrared-visible pedestrian re-recognition is improved, solving the shortcomings of feature alignment and local feature learning in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113887353B_ABST
    Figure CN113887353B_ABST
Patent Text Reader

Abstract

The present invention provides a visible light-infrared pedestrian re-identification method and system, which obtain an image to be identified; use a deep network model to extract features of an image pair to obtain local features and global features, use a posture estimation network model to extract features of the image to obtain a heat map feature, perform a dot product operation on the local features and the heat map and then connect them with the global features to obtain image features; obtain graph features based on the acquired image features and a preset joint point adjacency matrix; obtain final features of the image to be identified based on the obtained graph features and a preset graph convolutional network, and obtain a pedestrian re-identification result based on the similarity between the final features of the image to be identified and the corresponding image features in a candidate image library; the present invention improves the discrimination ability of cross-modal pedestrian images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of cross-modal image recognition, and in particular to a visible light-infrared pedestrian re-identification method and system. Background Art

[0002] The statements in this section merely provide background art related to the present disclosure and do not necessarily constitute prior art.

[0003] In visible-infrared person re-identification (VI-ReID), due to the cross-modal differences between images taken by cameras with different spectra, many feature information such as color information cannot be used due to the change of modality, and even become an obstacle to accuracy in some methods. Therefore, feature extraction and learning of VI-ReID are extremely challenging. Most existing VI-ReID methods use deep learning methods to extract image features for direct matching, ignoring the potential relationship between features, which limits the performance of VI-ReID.

[0004] Visible-infrared pedestrian re-identification (VI-ReID) aims to find a designated person in cameras with different modalities that are not related to each other. This technology can be applied to urban management, public security prevention and control, and the construction of smart cities. Recently, it has attracted great research interest, but many research methods on VI-ReID are limited to one modality. The visible light pedestrian images taken by the camera mainly rely on the unique appearance features of pedestrians under visible light. However, at night or in dim conditions, the camera cannot capture clear visible light images of pedestrians. In order to ensure that the camera can work normally under insufficient light, most surveillance cameras can automatically switch from visible light (RGB) mode to near infrared (IR) mode to capture infrared images of pedestrians. However, there are some difficulties in the current research on visible-infrared pedestrian re-identification: the traditional RGB pedestrian image is an image taken in a visible light environment, which contains three channels of color information, while the IR image taken in the near infrared mode only contains one channel of invisible light information. In this case, the color information of the image cannot be used, so this cross-modal pedestrian matching is more difficult.

[0005] In some recent studies, some methods use methods similar to traditional target detection. These methods focus on the Euclidean distance of the difference between the two modalities and achieve feature alignment by directly minimizing the Euclidean distance of the two modal images. However, due to the huge color difference between the two modal images, this method is difficult to achieve the desired effect. Inspired by the traditional single-modal pedestrian re-identification, some methods use neural networks to learn the global features shared by different modalities, but this ignores the differences within the modalities, which may cause mismatches in some images. The use of part-based methods requires finding effective local features in two images with large modal differences. When there is a large difference between the two images, the local features learned usually contain more invalid information, which is easy to cause mismatches in the matching stage. Recently, there are also methods inspired by Generative Adversarial Networks (GAN) networks. The GAN is used to generate images to convert images of different modalities into images of the same modality for processing. However, in the pedestrian re-identification task, the labels of the training set and the test set are not shared, so the trained GAN network may not generate satisfactory images in the test stage. Summary of the invention

[0006] In order to address the deficiencies of the prior art, the present disclosure provides a visible light-infrared pedestrian re-identification method and system, which improves the discrimination capability of cross-modal pedestrian images.

[0007] In order to achieve the above objectives, the present disclosure adopts the following technical solutions:

[0008] A first aspect of the present invention provides a visible light-infrared pedestrian re-identification method.

[0009] A visible light-infrared pedestrian re-identification method includes the following processes:

[0010] Obtain an image to be recognized;

[0011] The deep network model is used to extract features from the image pair to obtain local features and global features. The posture estimation network model is used to extract features from the image to obtain heat map features. The local features and the heat map are dot-producted and then connected with the global features to obtain image features.

[0012] According to the acquired image features and the preset joint point adjacency matrix, the graph features are obtained;

[0013] According to the obtained graph features and the preset graph convolutional network, the final features of the image to be identified are obtained, and the pedestrian re-identification result is obtained according to the similarity between the final features of the image to be identified and the corresponding image features in the candidate image library;

[0014] The candidate image library includes at least visible light images and infrared images, and both the visible light images and the infrared images in the candidate image library have undergone the same feature extraction process as the image to be identified.

[0015] Furthermore, each layer of the graph convolutional network has a branching structure, each layer is connected to all previous layers, and the output of all previous layers is aggregated in the output of the last layer.

[0016] Furthermore, each branch is equipped with an attention weight, which dynamically assigns different weights to different convolutional layer outputs. The attention mechanism is used to change the weight of each layer, and the attention weight is incorporated into the model training. The weight is iteratively updated through back propagation, and the similarity measurement result is passed into the loss function as the basis for back propagation.

[0017] Furthermore, the output representation of each layer of the graph convolutional network is obtained;

[0018] According to the obtained output representation, the semantic features of each node at different layers are obtained;

[0019] Based on the semantic features of each node at different layers, the features of all nodes at all layers are obtained;

[0020] A first similarity measurement result is obtained according to the Euclidean distance between the features of all nodes in all layers and the corresponding features of the original image.

[0021] Furthermore, the graph convolutional network includes a dynamic feature update layer, including:

[0022] Get the features of all nodes in all layers;

[0023] Process the input semantic features and output a dynamic feature matrix;

[0024] After all nodes in all layers have their features multiplied by the dynamic feature matrix, the final semantic representation matrix is ​​obtained;

[0025] Obtaining a second similarity measurement result according to the cosine similarity between the semantic representation matrix and the semantic representation matrix corresponding to the original image;

[0026] A final similarity score is obtained according to the first similarity measurement result and the second similarity measurement result.

[0027] Furthermore, the joint point adjacency matrix includes:

[0028] When some parts of the body are connected, the node value of the two parts in the corresponding adjacency matrix is ​​1. If there is no connection, the node value is 0. The last column of the joint adjacency matrix is ​​used to represent the global features, and the values ​​of the last column are all 1.

[0029] Furthermore, if the acquired image to be identified is a visible light image, it is converted into a grayscale image.

[0030] Furthermore, both the deep network model and the posture estimation network model use triplet loss and cross entropy loss to update the model parameters.

[0031] A second aspect of the present invention provides a visible light-infrared pedestrian re-identification system.

[0032] A visible light-infrared pedestrian re-identification system, comprising:

[0033] The data acquisition module is configured to: acquire an image to be identified;

[0034] The image feature extraction module is configured to: use the deep network model to extract features from the image pair to obtain local features and global features, use the posture estimation network model to extract features from the image to obtain heat map features, perform dot product operations on the local features and the heat map, and then connect them with the global features to obtain image features;

[0035] A graph feature acquisition module is configured to: obtain graph features according to the acquired image features and a preset joint point adjacency matrix;

[0036] The re-identification module is configured to obtain the final features of the image to be identified based on the obtained graph features and the preset graph convolutional network, and obtain the pedestrian re-identification result based on the similarity between the final features of the image to be identified and the corresponding image features in the candidate image library.

[0037] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the visible light-infrared pedestrian re-identification method as described in the first aspect of the present invention.

[0038] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the visible light-infrared pedestrian re-identification method as described in the first aspect of the present invention are implemented.

[0039] Compared with the prior art, the beneficial effects of the present invention are:

[0040] 1. The method, system, medium or electronic device described in the present invention includes a feature extraction stage, a feature learning stage and a matching update stage. Compared with previous methods, the recognition accuracy of infrared-visible light pedestrian re-identification can be further improved.

[0041] 2. The method, system, medium or electronic device described in the present invention innovatively provides a new graph convolutional neural network that can solve the degradation problem existing in traditional neural networks. It connects the convolutional layers through a branch structure and uses the attention mechanism to dynamically change the weight of each layer to learn the hidden feature information in the network.

[0042] 3. The method, system, medium or electronic device described in the present invention proposes a dynamic feature update layer, which can highlight discriminative features and suppress meaningless features.

[0043] Advantages of additional aspects of the present disclosure will be given in part in the following description and in part will become apparent from the following description or will be learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.

[0045] Figure 1 A brief flowchart of the visible light-infrared pedestrian re-identification method based on key point feature extraction provided in Example 1 of the present invention.

[0046] Figure 2 A schematic diagram of the overall process of a visible light-infrared pedestrian re-identification algorithm based on key point feature extraction provided in Example 1 of the present invention;

[0047] Figure 3 Schematic diagram of the feature learning network (graph convolutional network and dynamic feature update layer) provided in Example 1 of the present invention. DETAILED DESCRIPTION

[0048] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0049] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.

[0050] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0051] In the absence of conflict, the embodiments in the present disclosure and the features in the embodiments may be combined with each other.

[0052] Embodiment 1:

[0053] Embodiment 1 of the present disclosure provides a visible light-infrared pedestrian re-identification method based on key point feature extraction, including three parts: feature extraction, feature learning and matching update.

[0054] In the feature extraction part, two independent networks are used to extract the features and location information of the image respectively and then concatenate the two to convert them into graph features;

[0055] In the feature learning part, the new graph convolution structure and branch structure are used to consider the valuable information of the intermediate layer and pass it to the matching and updating layer as the final output, automatically highlighting the discriminative features and suppressing meaningless features;

[0056] In the matching and updating part, the features obtained in the feature learning part are used to measure the similarity, and the features and similarities are passed into the loss function to update the network parameters. After a certain number of iterative training, the network achieves the best effect.

[0057] Specifically, Figure 1 and Figure 2 As shown in FIG. 1 , the method includes: constructing a pedestrian feature extraction network mainly composed of a ResNet50 feature extraction network and a HR-Net posture estimation network and performing pre-training; then merging the extracted features with a predefined joint point adjacency matrix to obtain processed graph features; and evaluating the network parameters using a loss function (Formula (1)), which is specifically a combination of a cross entropy loss and a triplet loss. The network can accurately locate the key point positions of pedestrians and extract feature vectors containing rich information.

[0058] After the feature extraction network, a feature learning network is constructed. The network uses an improved graph convolution network to learn the incoming graph features. Branch structures and attention parameters are added so that the network can take into account all the learned effective information. A matching and update layer is added after the graph convolution to purposefully filter the features learned by the graph convolution, enhance effective information, and suppress invalid information. This part will perform similarity matching on the input image, and the similarity result obtained is the predicted value. At the same time, the similarity score will be passed into the loss function as the basis for network update, and it will be updated iteratively through back propagation.

[0059] The above network needs to be trained and optimized before use, specifically:

[0060] The training process mainly includes two stages, namely the training stage and the testing stage. The training stage mainly uses the training set in the data set, while the testing stage mainly uses the pictures in the gallery set. The role of the training stage is to enable the network to learn which parts of the picture have discriminative information, so that the network can update parameters to learn discriminative features. The testing stage is used to evaluate the robustness of the trained network and obtain a preliminary training accuracy. Since the pictures used in the testing stage are shared with the training set pictures, the accuracy obtained in the testing stage is only for reference and is not used as the final evaluation standard.

[0061] After a round of overall network training, the network will undergo a performance evaluation using test set images. The purpose is to match the test set images with the gallery set images. The test set images are not shared with the training set images, so the robustness of the network to image processing can be tested. The evaluation results obtained at this time are an important indicator for measuring the performance of a network.

[0062] The two sub-networks of the overall network and a module responsible for prediction and update (feature extraction network, feature learning network and matching and update module) contain the loss functions corresponding to the two sub-networks in the matching and update module. The sub-network parameters are independent. The matching and update module uses back propagation to update the network parameters of the sub-networks and continuously trains until the network converges.

[0063] The training set images are passed in, and the predicted classification results are obtained after three parts of processing. The features and similarity scores are passed in the loss function, the loss values ​​of the two sub-networks are calculated and the sub-networks are back-propagated to update the parameters in the sub-networks. After the training is completed, the network is evaluated. By passing the test set images into the network model, a more accurate classification result can be obtained, achieving good visible light-infrared pedestrian re-identification effect.

[0064] In the specific implementation example, the network construction process is:

[0065] The network proposed in this embodiment includes two sub-networks: a feature extraction network and a feature learning network.

[0066] For the feature extraction network, given an input image, the image is first enlarged to 256 × 128, and the image is horizontally and randomly inverted. In the CNN backbone, ResNet50 pre-trained on the COCO dataset is used as the CNN backbone, and HR-Net is used as the pose estimation network, which is also pre-trained on the COCO dataset.

[0067] The number and categories of key points of the human body are obtained through manual division in advance and stored in an adjacency matrix. It is stipulated that the human body is divided into 14 key points by the network, and the corresponding graph also contains 14 nodes. Then the information output by ResNet50 and HR-Net is multiplied and connected with the adjacency matrix to obtain the graph features, and then the graph features are passed into the loss function to obtain the loss value. The loss function is composed of a cross entropy function and a triple function.

[0068] The feature learning network adds a branch structure to each layer based on the graph convolutional network, so that each layer is connected to all previous layers, and adds an attention parameter to the branch. The parameter changes continuously with training, and the output of all previous layers is brought together in the output of the last layer, so that the network can obtain the key information learned by the intermediate layers that may be ignored by traditional graph convolution.

[0069] After obtaining the learned features, in order to highlight the discriminative features, a dynamic feature update layer is added. A series of convolution operations are used to dynamically change different regions of the features. The feature regions with discriminative information are strengthened, while the regions with weak discriminative information or containing erroneous information are suppressed. At the same time, a similarity measure is output after the graph convolution network and the dynamic feature update layer. The two are used together as the basis for similarity discrimination. At the same time, the similarity measure is also passed into the loss function as the basis for back propagation. Both sub-networks use the Adam optimizer to optimize the network, with a learning rate of 0.00035 and a weight decay of 0.0005. The training and testing codes are written in Pytorch and run on NVIDIA Titan RTX 24GB.

[0070] In a more specific implementation example, the goal of the feature extraction network is to process the image to obtain features that can be used for subsequent network learning. Combined with the needs of the feature learning network, the feature extraction network contains two sub-networks for extracting features and obtaining position information respectively. Among them, local features and global features are extracted by the ResNet50 convolutional network, while position information is extracted by the HR-Net convolutional network. Since the information at this time is an independent individual, in order to integrate them into an interconnected whole, an adjacency matrix is ​​specially constructed to save the human joint structure information, which is combined with the features to turn the scattered information obtained by the two sub-networks into a whole information. At this time, the information becomes a graph feature that can be learned by the graph convolutional network.

[0071] In visible light-infrared pedestrian re-identification, the infrared image in the identified image does not have color information, so the color information has no discrimination ability in the method described in this embodiment and may even inhibit the discrimination effect. Therefore, before the image is transmitted to the network, the image is uniformly converted into black and white images.

[0072] For the ResNet50 network, local features F can be obtained by processing the image l With the global feature F g , and the HR-Net network can obtain the heat map H of the image. The heat map can focus the line of sight on the key points of pedestrians, making other insignificant positions such as background and obstacles occupy a smaller proportion. Then, the local features and the heat map are subjected to dot product operation, so that the value representing the key position in the feature has a larger proportion, and the feature of its position is suppressed. At the same time, in order to consider the global information of the image, it is directly connected with the global feature after the dot product operation. Finally, the feature information combined with the heat map is obtained through global average pooling, as shown in formula (1).

[0073]

[0074] Where a represents the image category. When a is equal to IR, it means the image is an infrared image. When a is equal to RGB, it represents a visible light image. ⊙ represents the dot product operation of the vector. g[·] represents global average pooling. [,] represents the concatenation of two vectors.

[0075] The new feature matrix F a There is no direct connection between the internal features. According to the relationship between different parts of the human body, this embodiment defines an adjacency matrix A to save the relationship between features. When some parts of the body are connected (such as the left shoulder and the left elbow), the node value of the two parts in the corresponding adjacency matrix is ​​1. If there is no connection, the node value is 0. Since the output of the network contains global features, an additional column is added to the end of the adjacency matrix to represent the global features. The last column value is all 1. The adjacency matrix can be used to connect unrelated features.

[0076] The two sub-networks use triplet loss and cross entropy loss to update the model parameters. The triplet loss is used to enhance the network’s ability to distinguish details of pedestrian images of different modalities; the cross entropy loss is used to speed up the network update and enable the network to learn stable features, as shown in equations (2)-(4).

[0077]

[0078] L cls = -logF a (3)

[0079]

[0080] Among them, I represents the number of key points, α i , represents the confidence of the i-th key point. In particular, the I+1-th position represents the global feature, Fa ai 、F a pi 、F a ni Represent anchor point features, positive sample features and negative sample features respectively, || 2 Represents the Euclidean distance metric.

[0081] Some previous studies have proven the effectiveness of graph convolutional networks in pedestrian re-identification tasks, but traditional graph convolutional networks have some problems. For example, traditional graph convolutional networks are executed sequentially from front to back, and the output values ​​of each layer are only passed to the next layer in sequence, which may cause some useful information generated by the intermediate layer to be lost during the transmission process. Therefore, in this embodiment, on the basis of the traditional graph convolutional network, a residual structure similar to that in ResNet50 is added to solve the possible network degradation problem. However, unlike ResNet50, the branch structure of this embodiment connects each convolutional layer to each other, rather than the jump structure in ResNet50, and each connection carries an attention weight for dynamically giving different weights to different convolutional layer outputs.

[0082] Specifically, this embodiment proposes a new graph convolution structure for learning useful semantic information that may be learned by each layer in the network, and adds an additional dynamic feature update layer at the end of the network to train the model so that the model pays more attention to the meaningful part of the semantic information.

[0083] The traditional graph convolutional network is shown in formula (5):

[0084] H l+1 =σ(AH l W l ) (5)

[0085] Among them, H l+1 is the output of the l+1th layer, A represents the adjacency matrix of the relationship between nodes, H l represents the output of a layer on the network. In particular, H 0 is the graph feature output by the feature extraction network, W l is the weight matrix of the lth layer, and σ(·) is the activation function.

[0086] The graph convolutional network proposed in this embodiment uses the combination of local features and global features to form graph features as input, such as Figure 3As shown in the figure, due to the introduction of branches in the network, the network is no longer a simple serial operation. Each layer no longer only considers the output of the previous layer, but takes into account the output of all previous layers and the original features. However, the importance of the features obtained in each layer is different for different layers, and the result cannot be obtained by simply accumulating and summing each layer. Therefore, the attention mechanism is used to change the weight of each layer, and the attention weight is incorporated into the model training. The weight is iteratively updated through back propagation, and the output of each layer is expressed as formula (6).

[0087]

[0088] Correspondingly, the semantic features of each node at different layers can be expressed as formula (7):

[0089]

[0090] Among them, H vn i Represents the semantic features belonging to the vn node in the semantic features output by each layer.

[0091] Then the semantic features of each node at different layers can be expressed as formula (8):

[0092] F vn =[F 0 ,F 1 ,…,F l ] (8)

[0093] Therefore, the characteristics of all nodes in all layers can be expressed as formula (9):

[0094] F=[F1,F2,…,F n ] (9)

[0095] Using the obtained feature F, the original image can be measured for similarity. The similarity measurement is evaluated by calculating the Euclidean distance between the image pairs, as shown in formula (10).

[0096] S 1 x1,x2 =σ(f|F x1 -F x2 |) (10)

[0097] Among them, σ(·) is the activation function, f is a fully connected layer with the same dimension as the total number of pedestrian categories, and || represents the absolute value operation

[0098] In this embodiment, after the graph convolution outputs the learned feature information, a dynamic feature update layer is added so that the model can capture useful feature information more effectively. Specifically, the feature matrix obtained by formula (9) is passed into the dynamic feature update layer to obtain a feature representation matrix, which can focus on the discriminative feature information in the semantic features.

[0099] The dynamic feature update layer mainly includes two steps. First, the input semantic features are processed using formula (11) to output the dynamic feature matrix.

[0100] F OUT =σ(f2(tanh(f1(F T )))) (11)

[0101] Among them, σ represents the activation function, f1 and f2 represent two fully connected layers whose parameters are not shared, and F T It represents the transposed matrix of the feature matrix F. The dynamic update matrix obtained in formula (11) is represented as the weight of the graph convolution layer of different nodes. Then, the feature F is multiplied by the weight matrix to obtain the final semantic representation matrix, which is expressed as formula (12).

[0102] M=F OUT F (12)

[0103] The semantic representation matrix M is a feature matrix obtained by comprehensively considering the semantic information learned in each layer of the graph convolution and combining it with a dynamic update mechanism. Given two image pairs (x2, x2) of different modalities, the relevant features (M1, M2) are obtained through formula (12), so the cosine similarity can be calculated, see formula (13), where |||·||2 represents the second normal form.

[0104]

[0105] In the feature extraction network, cross entropy loss is used to optimize the entire model, and the input is the similarity obtained by formula (13), where y = 1 when the pedestrians in images x1 and x2 belong to the same category, otherwise y = 0.

[0106] L2=-[ylogS x1,x2 +(1-y)log(1-S x1,x2 )] (14)

[0107] Combining the feature extraction network and the feature learning network, the final loss function is obtained as shown in formula (15), where λ1 and λ2 are hyperparameters.

[0108] L=λ1L1+λ2L2 (15)

[0109] In the inference stage, by combining similarity formula (10) and formula (13), the given image x is compared with the gallery images through formula (16) to obtain the final similarity score, where δ is a hyperparameter.

[0110]

[0111] In a specific implementation example, the input image is first preprocessed, scaled to a specified size, and randomly inverted and cropped. The image is uniformly converted into a grayscale image and then sent to the network for training. The network parameters are updated using a back-propagation algorithm to achieve the best performance.

[0112] Specifically, the input image is first scaled in the program to make it uniformly 256×128, and the image is randomly horizontally inverted and randomly cropped, and the visible light image is converted to a grayscale image. In CNNbackbone, the ResNet50 network pre-trained with the COCO dataset is used as the CNN backbone, and the key point extraction network uses the currently more advanced HR-Net, which is also pre-trained on the COCO dataset. The number and category of human key points are obtained by artificial division in advance. Here, the network is specified to divide the human body into 14 key points, and the corresponding graph convolution network also contains 14 nodes. The algorithm is optimized using Adam, the learning rate is set to 0.00035, and the weight decay is 0.0005. The entire training and testing process is written in Pytorch and runs on NVIDIA Titan RTX 24GB.

[0113] In this embodiment, experiments are conducted on three datasets: RegDB, SYSU-MM01, and Market-1501.

[0114] RegDB is a VI-ReID dataset obtained by taking pictures of 412 different pedestrians from the front, back and side by a visible light camera and a near infrared camera. The dataset contains 8240 pictures in total, of which each person has ten RGB images of 100×110×3 pixels and ten IR images of 100×110×1 pixels. The method of this embodiment is compared with other existing methods. The comparison process uses two modes: visible-to-infrared and infrared-to-visible. The difference between the two modes lies in the difference in the identities of the two modes. In visible-to-infrared, visible images are used as query sets and thermal images are used as gallery sets, while infrared-to-visible is just the opposite.

[0115] SYSU-MM01 contains image information collected by six cameras (4 infrared cameras and 2 near-infrared cameras), which are divided into two cases: all-scene search and indoor scene search according to different search scenes. The training set and test set contain 395 categories and 96 categories respectively, totaling 491 categories. The data set is also divided into single-shot and multi-shot modes according to the number of training samples. There is only one target queried in single-shot, while there are 10 targets queried in multi-shot. The method of this embodiment is compared with the methods of others at present. In the comparison process, it is divided into four cases according to the different scenes and the number of samples.

[0116] Market-1501 contains 32668 images, including 32668 bounding boxes and 1501 identity categories. These images were taken by a total of 6 different and unrelated visible cameras. Of these images, 19732 images 230 (from 750 different people) form the training set.

[0117] The artificially assigned parameters included in this embodiment are as follows: in formula (15), the parameters λ1 and λ2 are assigned values ​​of 1 and 0.1 respectively; and in formula (16), the parameter δ is assigned a value of 0.5.

[0118] The algorithm implementation process is as follows:

[0119] Part 1: Feature Extraction Network

[0120] Input: dataset image x; dataset label;

[0121] Output: learned graph feature F

[0122] Initialization: Initialize the image size to 256×128, and convert the visible light image to a grayscale image; load the Resnet50 pre-trained model pre-trained on the ImageNet dataset and the HR-Net pre-trained model pre-trained on the COCO dataset into the corresponding models; set the batch size Batchsize = 14, and the number of human joints = 14.

[0123] Loop and execute the following statements

[0124] 1. Calculate the feature vector and heat map of the input image obtained through the pre-trained model

[0125] F l ,F g =ResNet50(x)

[0126] H = HR-Net(x)

[0127] 2. Merge the feature vector with the heat map

[0128]

[0129] 3. Convert features into graph features using the predefined joint adjacency matrix

[0130] F=HF

[0131] 4. Calculate the loss

[0132]

[0133] L cls = -logF a

[0134]

[0135] Stage 2: Feature Learning Network

[0136] Input: optimized feature information F; dataset image I; dataset label

[0137] Output: discriminant information and similarity scores from deep mining features

[0138] Initialization: Randomly initialize the parameters of each layer of the graph convolutional neural network, the attention parameter value, and randomly initialize the parameter value of the dynamic feature update layer

[0139] Loop and execute the following statements

[0140] 1. Pass graph features into graph convolutional neural network

[0141]

[0142] 2. Use the branch structure to combine each layer to get the output value

[0143]

[0144] F vn =[F 0 ,F 1 ,…,F l ]

[0145] F=[F1,F2,…,F n ]

[0146] 3. Use dynamic feature update mechanism to get the final output

[0147] F OUT =σ(f2(tanh(f1(F T ))))

[0148] 4. Calculate similarity and loss

[0149]

[0150] L2=-[ylogS x1,x2 +(1-y)log(1-S x1,x2 )]

[0151] 5. Combine the two losses of the two sub-networks and use the back propagation algorithm to update the network parameters

[0152] L=λ1L1+λ2L2

[0153] 6. In the matching stage, the first n nearest neighbors are obtained, and then the similarity of the first n neighbors is calculated. The neighbor with the highest similarity is taken as the prediction result.

[0154]

[0155] Experiments were conducted on three datasets, comparing eight other popular light-infrared pedestrian re-identification methods (HOG, LOMO, BDTR, Hi-CMD, CDP, D2RL, JSIA, HI-CMD) and three visible light pedestrian re-identification methods (PCB, PGFA, HOReID). It can be seen from Tables 1-3 that the method provided in this embodiment has better retrieval performance than other methods on different datasets.

[0156] Table 1: Comparison results on the SYSU-MM01 dataset

[0157]

[0158] Table 2: Comparison results on the RegDB dataset

[0159]

[0160] Table 3: Comparison results on the Market-1501 dataset

[0161] method R1 mAP PCB 92.3 77.4 PGFA 91.2 76.8 HOReID 94.2 84.9 The present invention 94.4 85.5

[0162] Embodiment 2:

[0163] Embodiment 2 of the present invention provides a visible light-infrared pedestrian re-identification system, including:

[0164] The data acquisition module is configured to: acquire an image to be identified;

[0165] The image feature extraction module is configured to: use the deep network model to extract features from the image pair to obtain local features and global features, use the posture estimation network model to extract features from the image to obtain heat map features, perform dot product operations on the local features and the heat map, and then connect them with the global features to obtain image features;

[0166] A graph feature acquisition module is configured to: obtain graph features according to the acquired image features and a preset joint point adjacency matrix;

[0167] The re-identification module is configured to obtain the final features of the image to be identified based on the obtained graph features and the preset graph convolutional network, and obtain the pedestrian re-identification result based on the similarity between the final features of the image to be identified and the corresponding image features in the candidate image library.

[0168] The working method of the system is the same as the visible light-infrared pedestrian re-identification method provided in Example 1, and will not be repeated here.

[0169] Embodiment 3:

[0170] Embodiment 3 of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the visible light-infrared pedestrian re-identification method as described in Embodiment 1 of the present invention.

[0171] Embodiment 4:

[0172] Embodiment 4 of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps in the visible light-infrared pedestrian re-identification method as described in Embodiment 1 of the present invention are implemented.

[0173] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0174] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0175] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0176] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0177] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0178] The above description is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A visible light-infrared pedestrian re-identification method, characterized by: The process includes: Obtain an image to be recognized; The deep network model is used to extract features from the image pair to obtain local features and global features. The posture estimation network model is used to extract features from the image to obtain heat map features. The local features and the heat map are dot-producted and then connected with the global features to obtain image features. According to the acquired image features and the preset joint adjacency matrix, the graph features are obtained; wherein the joint adjacency matrix includes: when some parts of the body are connected, the node value of the two parts in the corresponding adjacency matrix is ​​1, if there is no connection, then the node value is 0, the last column of the joint adjacency matrix is ​​used to represent the global features, and the values ​​of the last column are all 1; According to the obtained graph features and the preset graph convolutional network, the final features of the image to be identified are obtained, and the pedestrian re-identification result is obtained according to the similarity between the final features of the image to be identified and the corresponding image features in the candidate image library; each layer of the graph convolutional network has a branch structure, each layer is connected to all the previous layers, and the output of all previous layers is collected in the output of the last layer; each branch has an attention weight, which dynamically assigns different weights to different convolutional layer outputs, uses the attention mechanism to change the weight of each layer, incorporates the attention weight into the model training, and iteratively updates the weight through back propagation. The similarity measurement result is passed into the loss function as the basis for back propagation; Specifically, the output representation of each layer of the graph convolutional network is obtained; the semantic features of each node at different layers are obtained according to the obtained output representation; the features of all nodes at all layers are obtained according to the obtained semantic features of each node at different layers; the first similarity measurement result is obtained according to the Euclidean distance between the features of all nodes at all layers and the corresponding features of the original image; The graph convolutional network includes a dynamic feature update layer, including: obtaining features of all nodes in all layers; processing the input semantic features and outputting a dynamic feature matrix; obtaining a final semantic representation matrix after multiplying the features of all nodes in all layers with the dynamic feature matrix; obtaining a second similarity measurement result according to the cosine similarity between the semantic representation matrix and the semantic representation matrix corresponding to the original image; obtaining a final similarity score according to the first similarity measurement result and the second similarity measurement result; The candidate image library includes at least visible light images and infrared images.

2. The visible light-infrared pedestrian re-identification method according to claim 1, characterized in that: If the image to be identified is a visible light image, it is converted into a grayscale image; or, Both the deep network model and the pose estimation network model use triplet loss and cross entropy loss to update the model parameters.

3. A visible light-infrared pedestrian re-identification system, characterized by: include: The data acquisition module is configured to: acquire an image to be identified; The image feature extraction module is configured to: use the deep network model to extract features from the image pair to obtain local features and global features, use the posture estimation network model to extract features from the image to obtain heat map features, perform dot product operations on the local features and the heat map, and then connect them with the global features to obtain image features; The graph feature acquisition module is configured to: obtain graph features according to the acquired image features and the preset joint adjacency matrix; wherein the joint adjacency matrix includes: when some parts of the body are connected, the node value of the two parts in the corresponding adjacency matrix is ​​1, if there is no connection, then the node value is 0, and the last column of the joint adjacency matrix is ​​used to represent the global feature, and the value of the last column is all 1; The re-identification module is configured as follows: according to the obtained graph features and the preset graph convolution network, the final features of the image to be identified are obtained, and the pedestrian re-identification result is obtained according to the similarity between the final features of the image to be identified and the corresponding image features in the candidate image library; each layer of the graph convolution network is provided with a branch structure, each layer is connected to all the previous layers, and the output of all the previous layers is aggregated in the output of the last layer; each branch is provided with an attention weight, and different weights are dynamically assigned to different convolution layer outputs, and the weight of each layer is changed by using the attention mechanism, and the attention weight is incorporated into the model training, and the weight is iteratively updated through back propagation, and the similarity measurement result is passed into the loss function as the basis for back propagation; Specifically, the output representation of each layer of the graph convolutional network is obtained; the semantic features of each node at different layers are obtained according to the obtained output representation; the features of all nodes at all layers are obtained according to the obtained semantic features of each node at different layers; the first similarity measurement result is obtained according to the Euclidean distance between the features of all nodes at all layers and the corresponding features of the original image; The graph convolutional network includes a dynamic feature update layer, which includes: obtaining features of all nodes in all layers; processing the input semantic features and outputting a dynamic feature matrix; multiplying the features of all nodes in all layers with the dynamic feature matrix to obtain a final semantic representation matrix; obtaining a second similarity measurement result based on the cosine similarity between the semantic representation matrix and the semantic representation matrix corresponding to the original image; and obtaining a final similarity score based on the first similarity measurement result and the second similarity measurement result.

4. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, the steps in the visible light-infrared pedestrian re-identification method as described in any one of claims 1-2 are implemented.

5. An electronic device comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the visible light-infrared pedestrian re-identification method as described in any one of claims 1-2 are implemented.

Citation Information

Patent Citations

  • Behavior recognition method based on graph convolution and capsule neural network

    CN111461063A

  • Pedestrian re-identification method based on local graph convolutional network

    CN111695460A