A cross-modal pedestrian retrieval method for autonomous driving

By introducing local paired graph attention network and total contrast loss function in the cross-modal pedestrian search method, learning the relationship between pedestrian features in different modal pedestrians and optimizing model parameters, the problems of weakening of feature discriminantity and the optimization of modal feature distribution distance in the prior art are solved, and more efficient cross-modal pedestrian search performance is achieved.

CN115359554BActive Publication Date: 2025-05-23TIANJIN NORMAL UNIVERSITY +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210942178.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2025-05-23
Estimated Expiration
2042-08-08

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian search method ignores the use of pairs of feature relationships of different modal images to build graph nodes, and does not fuse additional information between nodes during the update process of graph nodes, resulting in weakening of feature discrimination; at the same time, the metric-based loss function ignores the relationship between a single modal feature and its heterogeneous center, resulting in the optimization of singular point features and modal feature distribution distances.

Method used

A cross-modal pedestrian retrieval method for autonomous driving is proposed. A local feature extraction module is constructed using pre-trained deep learning model, local features of visible light and infrared modal pedestrian images are obtained, and the relationship between different modal features is learned through local paired graph attention networks. At the same time, a total comparison loss function is constructed, combining local features and heterogeneous centers to optimize the cross-modal pedestrian retrieval model parameters.

Benefits of technology

By learning the relationship between pedestrian characteristics of different modalities, the differences between visible light and infrared pedestrian images are alleviated, and the performance and accuracy of cross-modal pedestrian retrieval is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359554B_ABST
    Figure CN115359554B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention discloses a cross-modal pedestrian retrieval method for autonomous driving, the method comprising: extracting local features of visible light modality training pedestrian images and infrared modality training pedestrian images; obtaining a local pairwise graph attention network corresponding to each region according to the difference between the local features; calculating heterogeneous centers according to the local features of the same identity and the same region, and constructing a total contrast loss function according to the local features and the corresponding heterogeneous centers; constructing a total loss function, and using the total loss function to update the parameters of the cross-modal pedestrian retrieval model to obtain an optimal cross-modal pedestrian retrieval model; and using the optimal cross-modal pedestrian retrieval model to perform cross-modal pedestrian retrieval. The present invention utilizes convolutional neural networks and graph attention mechanisms to learn the relationship between features of pedestrians trained in different modalities, ultimately alleviating the difference between visible light pedestrian images and infrared pedestrian images, and further improving the performance of cross-modal pedestrian retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision, pattern recognition, autonomous driving and artificial intelligence, and specifically relates to a cross-modal pedestrian retrieval method for autonomous driving. Background Art

[0002] Environmental perception is the first step in realizing intelligent driving, and it can solve the core problems of "where am I" and "where am I going" in autonomous driving. Cameras are important sensors for environmental perception. They can identify information such as geometric features and surface textures of objects, and detect obstacles through algorithms. The technology is mature and the cost is low. The maximum range is more than 500 meters. Image perception information in autonomous driving has received widespread attention. At the same time, pedestrians are one of the most important objects for analysis on the road.

[0003] In recent years, pedestrian retrieval has received extensive attention from academia and industry due to its important role in autonomous driving and intelligent monitoring. It mainly studies the method of searching for the same pedestrian under different cameras. However, ordinary pedestrian retrieval models cannot be applied to 24-hour all-weather monitoring. In order to overcome this limitation, cross-modal pedestrian retrieval technology has emerged. Cross-modal pedestrian retrieval not only faces the challenges of ordinary pedestrian retrieval, but also faces the huge challenge of modal differences.

[0004] Recently, many cross-modal pedestrian retrieval methods have proposed to overcome the modality difference by learning the feature expression of two modal images and using metric functions to reduce the distance between the two modal features. Some researchers have extracted specific modality features and shared modality features through deep models to better represent the two modal images. At the same time, using deep learning models to extract local features of two modal images has also been shown to improve the performance of cross-modal pedestrian retrieval methods. In addition, some methods use graph convolution to aggregate pedestrian features of other identities to improve the representation of features. Ye et al. used the pedestrian image features of two modalities after global pooling as graph nodes, used the image labels to establish an adjacency matrix and constructed a graph attention mechanism. Through the adjacency matrix, the graph attention coefficient integrates different modal features into each global feature, thereby improving the robustness of the feature. Zhang et al. used the local features of pedestrians as graph nodes, and the local features of each image formed a graph. The adjacency relationship between nodes and the graph coefficient were used to alleviate the changes within the modality. Regarding the metric function, Ye et al. used a weighted triplet loss function to alleviate the differences between modalities, shortening the distance between the features of pedestrians with the same modality and identity, and the distance between the features of pedestrians with the same identity in different modalities. Zhu et al. proposed a heterogeneous center loss, using a deep learning network to extract local features of pedestrian images in two modalities, and then calculated the mean features of the same identity in the same region in the two modalities. In the optimization process, the distance between the mean features of the same identity in different modalities becomes smaller, thereby reducing the distance between modalities.

[0005] The above graph-based cross-modal pedestrian retrieval methods ignore the use of paired feature relationships between different modal images to construct graph nodes. In the process of updating graph nodes, no additional information between nodes is integrated, which weakens the discriminability of features. The metric-based loss function ignores the relationship between a single modal feature and its heterogeneous center, resulting in singular point features, which affects the optimization of the modal feature distribution distance. Summary of the invention

[0006] The purpose of the present invention is to alleviate the modal differences between pedestrian images of different modalities. To this end, the present invention provides a cross-modal pedestrian retrieval method for autonomous driving.

[0007] In order to achieve the above object, the present invention proposes a cross-modal pedestrian retrieval method for autonomous driving, comprising the following steps:

[0008] Step S1, using the pre-trained deep learning model to build a local feature extraction module, inputting the visible light modality training pedestrian image and the infrared modality training pedestrian image into the local feature extraction module, and obtaining P local features for each training pedestrian image;

[0009] Step S2, obtaining a local pairwise graph attention network corresponding to each region according to the difference between the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image;

[0010] Step S3, according to the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity and the same area, the heterogeneous center of the corresponding identity and area of ​​the other modality training pedestrian image is calculated, and the total contrast loss function is constructed according to the local features and their corresponding heterogeneous centers;

[0011] Step S4, constructing a total loss function, and using the total loss function to update the parameters of the cross-modal pedestrian retrieval model to obtain an optimal cross-modal pedestrian retrieval model, wherein the cross-modal pedestrian retrieval model includes a local feature extraction module, a local paired graph attention network, and a classification fully connected layer;

[0012] Step S5, in the testing phase, the optimal cross-modal pedestrian retrieval model is used to extract local features of multiple regions of the pedestrian image to be queried, the local features of the multiple regions are connected in series as query features of the pedestrian image to be queried, and a search is performed in the cross-modal pedestrian library image based on the query features to obtain a cross-modal pedestrian retrieval result.

[0013] Optionally, the step S1 includes the following steps:

[0014] Step S11, constructing two pre-trained deep learning models respectively, and initializing parameters thereof to obtain the local feature extraction module;

[0015] Step S12, preprocessing the visible light modality training pedestrian images and infrared modality training pedestrian images in the training set;

[0016] Step S13, inputting the preprocessed visible light modality training pedestrian image and infrared modality training pedestrian image into the local feature extraction module respectively, and obtaining feature maps corresponding to the visible light modality training pedestrian image and infrared modality training pedestrian image respectively;

[0017] Step S14, performing average pooling based on the obtained feature map to obtain P local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image respectively.

[0018] Optionally, step S2 includes the following steps:

[0019] Step S21: subtract the local features of the pth region of the i-th visible light modality training pedestrian image and the j-th infrared modality training pedestrian image to obtain the local feature difference of the same region in different modalities. and using the local feature difference as the jth node in the local pairwise graph attention network of the pth region, wherein the number of the local pairwise graph attention networks is P;

[0020] Step S22, calculating the distance between each node in each local pairwise graph attention network, and constructing a distance-based variable coefficient, wherein the smaller the distance between nodes, the larger the corresponding variable coefficient, thereby enhancing the correlation between nodes;

[0021] Step S23, connecting the nodes in each local paired graph attention network in pairs in series, and using the learnable parameter vector in combination with the variable coefficient to calculate the semantic graph attention coefficient between the nodes in each local paired graph attention network;

[0022] Step S24, using the semantic graph attention coefficients between nodes in each local pairwise graph attention network to update the corresponding nodes to learn the correlation between nodes and the relationship between different modalities;

[0023] Step S25, using the learnable parameter matrix and the nonlinear function to perform a final update on the node.

[0024] Optionally, the variable coefficient constructed based on the distance between nodes in the local pairwise graph attention network of the p-th region is expressed as:

[0025]

[0026] Among them, β is an adjustable parameter, || || 2 Indicates the calculation of Euclidean distance, and They represent the nth node and the mth node in the local pairwise graph attention network of the pth region, respectively.

[0027] Optionally, the semantic graph attention coefficient between the nth node and the mth node in the local pairwise graph attention network of the pth region is expressed as:

[0028]

[0029] in, represents the u-th node in the local pairwise graph attention network of the p-th region, l() represents the LeakyReLU operation, “,” represents the concatenation of node features, q∈R 2D×1 is a learnable parameter vector, D represents the dimension of nodes in the local pairwise graph attention network, and U is the total number of nodes in the local pairwise graph attention network for the p-th region.

[0030] Optionally, the nth node in the pth region local pairwise graph attention network updated with the semantic graph attention coefficient is expressed as:

[0031]

[0032] Optionally, the node representation obtained after the final update of the node using the learnable parameter matrix and the nonlinear function is:

[0033]

[0034] Among them, w T ∈R 2D×2 is a learnable matrix, Represents the nonlinear operation ELU function.

[0035] Optionally, step S3 includes the following steps:

[0036] Step S31, respectively calculating the mean features of all local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity in the pth region as the heterogeneous center of the corresponding identity and region of the other modality training pedestrian image;

[0037] Step S32, based on the similarity between all local features in each region and the corresponding heterogeneous center, a total contrast loss function is constructed using contrast learning.

[0038] Optionally, the total contrast loss function is expressed as:

[0039]

[0040] in:

[0041]

[0042]

[0043] τ represents the scale adjustment coefficient, S represents the total number of identities of pedestrian images trained in infrared mode, yes The heterogeneous center yes Heterogeneous center.

[0044] Optionally, step S4 includes the following steps:

[0045] Step S41, constructing a binary cross entropy loss function based on the nodes of the local paired graph attention network and whether the paired local features related to the nodes in the local paired graph attention network are of the same identity;

[0046] Step S42, constructing a cross-modal contrast loss function based on heterogeneous centers;

[0047] Step S43, obtaining a multi-classification cross entropy loss function according to the local features of the images of different modalities and the pedestrian labels of the images of different modalities

[0048] Step S44: construct a heterogeneous center loss function based on the heterogeneous center features Step S45, based on the binary cross entropy loss function Cross-modal contrastive loss function based on heterogeneous center Multi-classification cross entropy loss function Heterogeneous center loss function A total loss function is constructed, and P total loss values ​​of pedestrian images trained in different modalities are calculated using the total loss function, and then a mean loss value of the P total loss values ​​is obtained. The cross-modal pedestrian retrieval model is optimized using the mean loss value to obtain the optimal cross-modal pedestrian retrieval model.

[0049] The beneficial effects of the present invention are as follows: the present invention utilizes convolutional neural networks and graph attention mechanisms to learn the relationship between pedestrian features trained in different modalities, ultimately alleviating the differences between visible light pedestrian images and infrared pedestrian images, and further improving the performance of cross-modal pedestrian retrieval.

[0050] It should be noted that this invention was funded by the National Natural Science Foundation of China Project No.62171321, the Tianjin Natural Science Foundation Key Project No.20JCZDJC00180 and the Tianjin Normal University Graduate Research Innovation Project No.2022KYCX032Z. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 This is a flowchart of a cross-modal pedestrian retrieval method for autonomous driving according to an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings. It should be understood that these descriptions are only exemplary and are not intended to limit the scope of the present invention. In addition, in the following description, the description of well-known structures and technologies is omitted to avoid unnecessary confusion of the concept of the present invention.

[0053] The present invention learns the local features of visible light modality and infrared modality pedestrian images through convolutional neural networks, and uses these local features to construct a local pairwise graph attention network and a cross-center based contrast learning method, thereby alleviating the difference in modal distribution, improving the representation ability and completeness of local features, and improving the accuracy of cross-modal pedestrian retrieval.

[0054] Figure 1 is a flow chart of a cross-modal pedestrian retrieval method for autonomous driving according to an embodiment of the present invention. Figure 1 As shown, the cross-modal pedestrian retrieval method for autonomous driving includes the following steps:

[0055] Step S1, using the pre-trained deep learning model to build a local feature extraction module, input the visible light modality training pedestrian image and the infrared modality training pedestrian image into the local feature extraction module, and obtain P local features for each training pedestrian image. The local feature of the pth region of the i-th visible light modality training pedestrian image is expressed as Similarly, the local feature of the pth region of the jth infrared modality training pedestrian image is expressed as Among them, the superscript V refers to the visible light mode, and the superscript I refers to the infrared mode;

[0056] Furthermore, the step S1 includes the following steps:

[0057] Step S11, constructing two pre-trained deep learning models respectively, and initializing parameters thereof to obtain the local feature extraction module;

[0058] In one embodiment of the present invention, the pre-trained deep learning model used can be selected as the model component before the last overall pooling in the ResNet-50 model, that is, the last global pooling layer and fully connected layer in the ResNet-50 model are removed, and the step size of the fourth block in the ResNet-50 model is changed from the original default 2 to 1, and a shared fully connected layer is added as a dimensionality reduction layer. In one example of the present invention, the shared fully connected layer contains 512 neurons.

[0059] Step S12, preprocessing the visible light modality training pedestrian images and infrared modality training pedestrian images in the training set;

[0060] In one embodiment of the present invention, preprocessing the visible light modality training pedestrian image and the infrared modality training pedestrian image includes: cropping the size of the training pedestrian image to a preset size, such as 288×144 pixels, and proportionally reducing all pixel values ​​of the training pedestrian image to a preset range, such as between 0 and 1, and then subtracting the pixel average of the training pedestrian image from each pixel value in the training pedestrian image, and then dividing the result by the pixel variance of the training pedestrian image.

[0061] Step S13, inputting the preprocessed visible light modality training pedestrian image and infrared modality training pedestrian image into the local feature extraction module respectively, and obtaining feature maps corresponding to the visible light modality training pedestrian image and infrared modality training pedestrian image respectively;

[0062] In one embodiment of the present invention, the number of the feature maps is 2048, and the size is 18×9, where 18 is the height of the feature map and 9 is the width of the feature map.

[0063] Step S14, performing average pooling based on the obtained feature map to obtain P local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image, respectively, wherein the local feature of the pth region of the i-th visible light modality training pedestrian image is expressed as The local feature of the pth region of the jth infrared modality training pedestrian image is expressed as Among them, the superscript V refers to the visible light mode, and the superscript I refers to the infrared mode;

[0064] In one example of the present invention, the shared fully connected layer contains 512 neurons. After the local features of the pth region of the i-th visible light modality training pedestrian image are reduced in dimension through the shared fully connected layer, it can be obtained: Similarly, after reducing the dimension of the local features of the pth region of the jth infrared modality training pedestrian image through the shared fully connected layer, we can obtain

[0065] For example, for the i-th visible light modality training pedestrian image, if P = 6, then after average pooling, 6 local features will be obtained: Then, the shared fully connected layer is used for dimensionality reduction, and 6 local features after dimensionality reduction can be obtained.

[0066] Step S2, obtaining a local pairwise graph attention network corresponding to each region according to the difference between the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image;

[0067] Furthermore, the step S2 comprises the following steps:

[0068] Step S21: subtract the local features of the pth region of the i-th visible light modality training pedestrian image and the j-th infrared modality training pedestrian image to obtain the local feature difference of the same region in different modalities. and using the local feature difference as the jth node in the local pairwise graph attention network of the pth region, wherein the number of the local pairwise graph attention networks is P;

[0069] In one embodiment of the present invention, the local feature difference of the same region in different modalities is calculated based on the local feature difference between the pedestrian images trained in different modalities, wherein the local feature difference of the same region in different modalities can be expressed as:

[0070]

[0071] in, and They respectively represent the local features of the p-th region of the i-th visible light modality training pedestrian image and the j-th infrared modality training pedestrian image.

[0072] Step S22, calculating the distance between each node in each local pairwise graph attention network, and constructing a distance-based variable coefficient, wherein the smaller the distance between nodes, the larger the corresponding variable coefficient, thereby enhancing the correlation between nodes;

[0073] In one example of the present invention, the variable coefficient constructed based on the distance between nodes in the local pairwise graph attention network of the p-th region can be expressed as:

[0074]

[0075] Among them, β is an adjustable parameter, || || 2 Indicates the calculation of Euclidean distance, and They represent the nth node and the mth node in the local pairwise graph attention network of the pth region, respectively.

[0076] It can be seen from the above formula that the smaller the distance between nodes, the larger the corresponding variable coefficient.

[0077] In one embodiment of the present invention, β=2.

[0078] Step S23, connecting the nodes in each local paired graph attention network in pairs in series, and using the learnable parameter vector in combination with the variable coefficient to calculate the semantic graph attention coefficient between the nodes in each local paired graph attention network;

[0079] In one example of the present invention, the semantic graph attention coefficient between the nth node and the mth node in the local pairwise graph attention network of the pth region is expressed as:

[0080]

[0081] in, represents the u-th node in the local pairwise graph attention network of the p-th region, l() represents the LeakyReLU operation, “,” represents the concatenation of node features, q∈R 2D×1 is a learnable parameter vector, D represents the dimension of nodes in the local pairwise graph attention network, and U is the total number of nodes in the local pairwise graph attention network for the p-th region.

[0082] In one embodiment of the present invention, q∈R 1024×1 , U=32.

[0083] Step S24, using the semantic graph attention coefficients between nodes in each local pairwise graph attention network to update the corresponding nodes to learn the correlation between nodes and the relationship between different modalities;

[0084] In one example of the present invention, the nth node in the pth region local pairwise graph attention network after being updated using the semantic graph attention coefficient is expressed as:

[0085]

[0086] Step S25, using the learnable parameter matrix and nonlinear function to perform a final update on the nodes to improve the representation ability of the nodes and the stability of the graph structure.

[0087] In one example of the present invention, the nodes updated by the semantic graph attention coefficient and the nodes before the update are connected in series, and the learnable parameter matrix and the nonlinear function are combined to obtain new nodes containing the difference relationship between different modalities and the distance information between nodes. The formula is as follows:

[0088]

[0089] Among them, w T ∈R 2D×2 is a learnable matrix, Represents the nonlinear operation ELU function.

[0090] In one embodiment of the present invention, wT ∈R 1024×2 .

[0091] Step S3, according to the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity and the same area, the heterogeneous center of the corresponding identity and area of ​​the other modality training pedestrian image is calculated, and the total contrast loss function is constructed according to the local features and their corresponding heterogeneous centers;

[0092] Furthermore, step S3 includes the following steps:

[0093] Step S31, respectively calculating the mean features of all local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity in the pth region as the heterogeneous center of the corresponding identity and region of the other modality training pedestrian image;

[0094] In an example of the present invention, the mean feature of the infrared modality training pedestrian image with identity b in the pth region is expressed as:

[0095]

[0096] Among them, O b represents the total number of local features of the pth region belonging to identity b in the infrared modality training pedestrian image, ID(i,I)=b means that the identity of the i-th infrared modality training pedestrian image is b, It represents the local features of the i-th infrared modality training pedestrian image in the p-th region, and the mean features of the infrared modality training pedestrian image with identity b in the p-th region. It can be used as the heterogeneous center of the corresponding area of ​​the pedestrian image trained in the visible light modality.

[0097] Similarly, the mean feature of the visible light modality training pedestrian image with identity b in the pth region can be calculated As the heterogeneous center of the corresponding region of the infrared modality training pedestrian image.

[0098] In one embodiment of the present invention, The dimension of O is 512. b =8.

[0099] Step S32, based on the similarity between all local features in each region and the corresponding heterogeneous centers, a loss function is constructed using contrastive learning to optimize the embedding space of the network using the loss function, supervise the network, improve the similarity between the local features and their corresponding heterogeneous centers, shorten the distance between the two, and further promote the local feature extraction module to extract more discriminative local features.

[0100] In an example of the present invention, a contrastive learning function is constructed using the similarity between a local feature and its corresponding heterogeneous center, wherein the contrastive learning function constructed based on the similarity between the local feature of the pth region of a visible light modality trained pedestrian image and its corresponding heterogeneous center can be expressed as:

[0101]

[0102] Among them, τ represents the scale adjustment coefficient, S represents the total number of identities of pedestrian images trained in infrared mode, yes Similarly, a contrastive learning function can be constructed based on the similarity between the local features of the pth region of the infrared modality trained pedestrian image and its corresponding heterogeneous center:

[0103]

[0104] in, yes Heterogeneous center.

[0105] Then, the total contrast loss function based on the heterogeneous center of the pth region can be obtained:

[0106] In an embodiment of the present invention, τ=0.2, S=4.

[0107] Step S4, constructing a total loss function, and using the total loss function to update the parameters of the cross-modal pedestrian retrieval model to obtain an optimal cross-modal pedestrian retrieval model, wherein the cross-modal pedestrian retrieval model includes a local feature extraction module, a local paired graph attention network, and a classification fully connected layer;

[0108] Furthermore, the step S4 comprises the following steps:

[0109] Step S41, constructing a binary cross entropy loss function based on the nodes of the local paired graph attention network and whether the paired local features related to the nodes in the local paired graph attention network are of the same identity;

[0110] In one example of the present invention, the constructed binary cross entropy loss function of the p-th region can be expressed as:

[0111]

[0112] in, is the label value, when When , it means that the local features of different modalities of the pairs of nodes in the j-th node in the local pairwise graph attention network that constitute the p-th region belong to the same identity. Otherwise,

[0113] Step S42, the total contrast loss function constructed in step S3 is as a cross-modal contrastive loss function based on heterogeneous centers;

[0114] Step S43, obtaining a multi-classification cross entropy loss function according to the local features of the images of different modalities and the pedestrian labels of the images of different modalities To calculate the multi-classification loss value;

[0115] Among them, the multi-classification cross entropy loss of the local features of the pth region of the training pedestrian image can be expressed as:

[0116]

[0117] Among them, C represents the total number of pedestrian categories in the training set, is the value of the true category of the local feature of the pth region of the training pedestrian image. When c is the true category, otherwise, The predicted value indicating that the local feature of the p-th region of the ith training pedestrian image belongs to category c is obtained by performing classification prediction on the local feature of the p-th region of the ith training pedestrian image through the corresponding classification fully connected layer.

[0118] In an example of the present invention, the heterogeneous center loss function formula can be expressed as:

[0119]

[0120] in, and They are the heterogeneous centers of the pth region under identity s, S=4.

[0121] Step S45, based on the binary cross entropy loss function Cross-modal contrastive loss function based on heterogeneous center Multi-classification cross entropy loss function Heterogeneous center loss function A total loss function is constructed, and P total loss values ​​of pedestrian images trained in different modalities are calculated using the total loss function, and then the mean loss value of the P total loss values ​​is obtained. The mean loss value is used to optimize a cross-modal pedestrian retrieval model composed of a local feature extraction module, a local paired graph attention network, and a classification fully connected layer to obtain the optimal cross-modal pedestrian retrieval model.

[0122] In one embodiment of the present invention, the total loss function can be expressed as:

[0123]

[0124] Among them, λ 1 , 2 and λ 3 is the weight coefficient, for example, λ 1 =0.5,λ 2 =0.4,λ 3 =0.5.

[0125] Step S5, in the testing phase, the optimal cross-modal pedestrian retrieval model is used to extract local features of multiple regions of the pedestrian image to be queried, the local features of the multiple regions are connected in series as query features of the pedestrian image to be queried, and a search is performed in the cross-modal pedestrian library image based on the query features to obtain a cross-modal pedestrian retrieval result.

[0126] Among them, the cosine distance between features can be used to represent the similarity between images, and then obtain cross-modal pedestrian retrieval results.

[0127] For example, if the pedestrian image to be queried is an infrared pedestrian image The local features of multiple regions extracted by using the optimal cross-modal pedestrian retrieval model are connected in series to obtain the infrared modality pedestrian image. Query features Assume that the concatenated feature representation of the cross-modal pedestrian library image obtained by using the optimal cross-modal pedestrian retrieval model is: Then the similarity Dist between the images can be obtained based on the cosine distance between the two features, that is;

[0128]

[0129] Among them, cosdist(,) means calculating the cosine distance.

[0130] The image in the cross-modal pedestrian database that has the highest similarity to the pedestrian image in the to-be-queried modality is the cross-modal pedestrian retrieval result of the pedestrian image in the to-be-queried modality.

[0131] The cross-modal pedestrian retrieval database publicly available on the Internet is used as the test object. For example, when tested on the single-shot and all-search modes of the SYSU-MM01 database, the correct search rate of the cross-modal pedestrian retrieval of the present invention reaches 61.89% (rank-1) and 60.12% (mAP). The cross-modal pedestrian retrieval method of the present invention uses the local pairwise graph attention network to effectively learn the difference relationship between pedestrian images of different modalities, and uses contrastive learning to shorten the distance between heterogeneous centers and local features in the embedding space, thereby improving the discriminability of local features and greatly improving the search rate of cross-modal pedestrian retrieval. This shows the effectiveness of the method of the present invention.

[0132] It should be understood that the above specific embodiments of the present invention are only used to illustrate or explain the principles of the present invention, and do not constitute a limitation of the present invention. Therefore, any modifications, equivalent substitutions, improvements, etc. made without departing from the spirit and scope of the present invention should be included in the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or the equivalent forms of such scope and boundaries.

Claims

1. A cross-modal pedestrian retrieval method for autonomous driving. It is characterized in that The method comprises the following steps: Step S1, using the pre-trained deep learning model to build a local feature extraction module, inputting the visible light modality training pedestrian image and the infrared modality training pedestrian image into the local feature extraction module, and obtaining P local features for each training pedestrian image; Step S2, obtaining a local pairwise graph attention network corresponding to each region according to the difference between the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image; Step S3, according to the local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity and the same area, the heterogeneous center of the corresponding identity and area of ​​the other modality training pedestrian image is calculated, and the total contrast loss function is constructed according to the local features and their corresponding heterogeneous centers; Step S4, constructing a total loss function, and using the total loss function to update the parameters of the cross-modal pedestrian retrieval model to obtain an optimal cross-modal pedestrian retrieval model, wherein the cross-modal pedestrian retrieval model includes a local feature extraction module, a local paired graph attention network, and a classification fully connected layer; Step S5, in the testing phase, the optimal cross-modal pedestrian retrieval model is used to extract local features of multiple regions of the pedestrian image to be queried, the local features of the multiple regions are connected in series as query features of the pedestrian image to be queried, and a search is performed in the cross-modal pedestrian library image based on the query features to obtain a cross-modal pedestrian retrieval result.

2. The method according to claim 1, It is characterized in that The step S1 comprises the following steps: Step S11, constructing two pre-trained deep learning models respectively, and initializing parameters thereof to obtain the local feature extraction module; Step S12, preprocessing the visible light modality training pedestrian images and infrared modality training pedestrian images in the training set; Step S13, inputting the preprocessed visible light modality training pedestrian image and infrared modality training pedestrian image into the local feature extraction module respectively, and obtaining feature maps corresponding to the visible light modality training pedestrian image and infrared modality training pedestrian image respectively; Step S14, performing average pooling based on the obtained feature map to obtain P local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image respectively.

3. The method according to claim 1, It is characterized in that The step S2 comprises the following steps: Step S21: subtract the local features of the pth region of the i-th visible light modality training pedestrian image and the j-th infrared modality training pedestrian image to obtain the local feature difference of the same region in different modalities. and using the local feature difference as the jth node in the local pairwise graph attention network of the pth region, wherein the number of the local pairwise graph attention networks is P; Step S22, calculating the distance between each node in each local pairwise graph attention network, and constructing a distance-based variable coefficient, wherein the smaller the distance between nodes, the larger the corresponding variable coefficient, thereby enhancing the correlation between nodes; Step S23, connecting the nodes in each local paired graph attention network in pairs in series, and using the learnable parameter vector in combination with the variable coefficient to calculate the semantic graph attention coefficient between the nodes in each local paired graph attention network; Step S24, using the semantic graph attention coefficients between nodes in each local pairwise graph attention network to update the corresponding nodes to learn the correlation between nodes and the relationship between different modalities; Step S25, using the learnable parameter matrix and the nonlinear function to perform a final update on the node.

4. The method according to claim 3, It is characterized in that The variable coefficient constructed based on the distance between nodes in the local pairwise graph attention network of the p-th region is expressed as: Among them, β is an adjustable parameter, || || 2 Indicates the calculation of Euclidean distance, and They represent the nth node and the mth node in the local pairwise graph attention network of the pth region, respectively.

5. The method according to claim 3, It is characterized in that The semantic graph attention coefficient between the nth node and the mth node in the local pairwise graph attention network of the pth region is expressed as: in, represents the u-th node in the local pairwise graph attention network of the p-th region, l( ) represents the LeakyReLU operation, Indicates that the node features are connected in series, q∈R 2D×1 is a learnable parameter vector, D represents the dimension of nodes in the local pairwise graph attention network, and U is the total number of nodes in the local pairwise graph attention network for the p-th region.

6. The method according to claim 3, It is characterized in that The nth node in the pth region local pairwise graph attention network after being updated with the semantic graph attention coefficient is expressed as:

7. The method according to claim 3, It is characterized in that The node representation obtained after the final update of the node using the learnable parameter matrix and nonlinear function is: Among them, w T ∈R 2D×2 is a learnable matrix, Represents the nonlinear operation ELU function.

8. The method according to claim 1, It is characterized in that The step S3 comprises the following steps: Step S31, respectively calculating the mean features of all local features of the visible light modality training pedestrian image and the infrared modality training pedestrian image of the same identity in the pth region as the heterogeneous center of the corresponding identity and region of the other modality training pedestrian image; Step S32, based on the similarity between all local features in each region and the corresponding heterogeneous center, a total contrast loss function is constructed using contrast learning.

9. The method according to claim 8, It is characterized in that The total contrast loss function is expressed as: in: τ represents the scale adjustment coefficient, S represents the total number of identities of pedestrian images trained in infrared mode, yes The heterogeneous center yes Heterogeneous center.

10. The method according to claim 1, It is characterized in that The step S4 comprises the following steps: Step S41, constructing a binary cross entropy loss function based on the nodes of the local paired graph attention network and whether the paired local features related to the nodes in the local paired graph attention network are of the same identity; Step S42, constructing a cross-modal contrast loss function based on heterogeneous centers; Step S43: Obtain the multi-class cross-entropy loss function based on the local features of different modality images and the pedestrian labels of different modality images Step S44: construct a heterogeneous center loss function based on the heterogeneous center features Step S45, based on the binary cross entropy loss function Cross-modal contrastive loss function based on heterogeneous center Multi-classification cross entropy loss function Heterogeneous center loss function A total loss function is constructed, and P total loss values ​​of pedestrian images trained in different modalities are calculated using the total loss function, and then a mean loss value of the P total loss values ​​is obtained. The cross-modal pedestrian retrieval model is optimized using the mean loss value to obtain the optimal cross-modal pedestrian retrieval model.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on local information learning

    CN112434796A

  • Cross-modal neural network construction method, pedestrian retrieval method and system

    CN113743544A