A Cross-modal Person Re-identification Method Based on Coal Mine Scenes
By constructing a cross-modal pedestrian re-identification method based on relational modeling and spectrum transformation in coal mine scenarios, the problems of insufficient samples, easy loss of feature information and low image quality in the prior art are solved, and more efficient pedestrian image feature representation and cross-modal retrieval performance are achieved.
Patent Information
- Application Number
- CN202411828573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-12
AI Technical Summary
The existing cross-modal pedestrian re-identification technology faces problems such as insufficient samples, easy loss of feature information and low image quality in coal mine scenarios, resulting in poor application effect in low light, high temperature and smoke environments.
A cross-modal pedestrian re-identification method based on coal mine scenes is proposed. By acquiring coal mine infrared light images and visible light images, using text pre-trained models to supplement missing color information, and a cross-modal pedestrian re-identification backbone network based on relational modeling and spectrum transformation is constructed, including color encoder, text encoder, fragment relationship modeling module and channel spectrum transformation module, enhancing the representation ability of pedestrian image features.
Effectively fusion of image features across modalities improves the richness and accuracy of pedestrian image features and improves the model's prediction and generalization ability in cross-modal pedestrian image retrieval scenarios.
Smart Images

Figure CN119785380B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of cross-modal person re-identification, and more specifically, relates to a cross-modal person re-identification method based on a coal mine scenario. Background Art
[0002] With the rapid development of economy and technology, it is becoming increasingly important to ensure personal and property safety in public places. The government and relevant departments have increased their investment in intelligent security. More and more cameras are put into use in public places, and intelligent monitoring has penetrated into people's lives and become an important tool for dealing with security problems and emergency events. In this context, the person re-identification (Re-ID) technology has attracted extensive attention from scholars and has achieved rapid development. Most of the existing person re-identification methods are dedicated to retrieving visible light images captured by cameras during the day. However, these methods cannot effectively retrieve pedestrian information in the low-light environment, especially in the dark environment of coal mine scenarios, which limits their effective practical applications. Therefore, cross-modal person re-identification has emerged. In recent years, the methods proposed by researchers to solve the cross-modal person re-identification problem can be mainly divided into two categories: First, the method based on feature sharing focuses on learning the shared features of pedestrians in cross-modal images to narrow the differences between modalities; second, the method based on feature complementation aims to achieve mutual complementation by learning the features of different parts.
[0003] However, the cross-modal person re-identification in the prior art has the following defects: First, the existing cross-modal datasets usually have obvious deficiencies in the diversity and richness of samples. The number and coverage of pedestrian samples are far less than those of single-modal datasets. Especially in the coal mine environment, environmental changes caused by factors such as light changes, personnel clothing, dust, and smoke further limit the representativeness and diversity of the dataset, and the generalization ability is greatly restricted. Due to the widespread existence of environmental factors such as low light, high temperature, and smoke in coal mine operation sites, infrared imaging is greatly affected, often with poor resolution and relatively blurred image details. The low-quality images not only increase the difficulty of feature extraction but also easily introduce noise and artifacts, further affecting the training effect of the model.
[0004] Second, the existing person re-identification methods based on feature sharing face serious problems of loss of feature information and a large amount of loss of effective frequency information in complex environments such as coal mines. Specifically: The obtained feature information is extremely vulnerable to factors such as the posture changes of workers in coal mine scenarios and noise interference in coal mine operation sites, resulting in the model being difficult to access sufficient diversity information such as identity, posture, perspective, and environmental changes during the training process, directly leading to the loss of important image feature information, especially the loss of a large amount of effective frequency information.
[0005] Thirdly, existing pedestrian re-identification methods based on feature supplementation face many challenges in applications in special environments such as coal mines. There is a lack of sufficient interaction between pedestrian samples and it is difficult to extract more effective pedestrian features. Specifically, affected by factors such as the low contrast of infrared images and insufficient texture information in the coal mine environment, the supplementary features often cannot fully reflect the true identity information of pedestrians. The model faces great difficulties in extracting effective pedestrian features and is difficult to handle the differences between heterogeneous information. Moreover, due to the complex environment in the coal mine operation area, factors such as soot and insufficient lighting, and the feature supplementation method mainly focuses on the feature enhancement of individual samples and lacks in-depth exploration of the relationships between samples, resulting in a lack of sufficient interaction and information sharing between pedestrian samples, restricting the accurate judgment of the pedestrian identity by the model and affecting its generalization ability and robustness in complex scenarios. Summary of the Invention
[0006] In order to solve the problems in the prior art such as the insufficient richness of pedestrian image dataset samples, the easy loss of pedestrian feature information, and the generally low quality of pedestrian images, the present invention provides a cross-modal pedestrian re-identification method based on a coal mine scenario, which effectively fuses the image features between cross-modalities and effectively mines the feature information of pedestrian images.
[0007] To achieve the above object, the present invention is implemented through the following technical solutions:
[0008] The present invention is a cross-modal pedestrian re-identification method based on a coal mine scenario, specifically including the following steps:
[0009] Step 1: Obtain the identity representing the pedestrian image and the identity label corresponding to the pedestrian image from the coal mine infrared light image and the coal mine visible light image. Each pedestrian image of different modalities includes a visible light image and an infrared light image, and preprocess the visible light image and the infrared light image. Pass the visible light image and the infrared light image through a text pre-training model for preprocessing to obtain the visible light image features and infrared light image features after the text pre-training model supplements the missing color information.
[0010] Step 2: Construct a cross-modal pedestrian re-identification backbone network based on relationship modeling and spectral transformation. The cross-modal pedestrian re-identification backbone network mainly includes: a color encoder, a text encoder, a segment relationship modeling module, a channel spectral transformation module, a visible light texture encoder, a joint relationship encoder, and an infrared texture encoder. Among them, the segment relationship modeling module includes a visual region image modeling module and a text region semantic modeling module, and a global-local cross-attention mechanism is added to each part of the visual region image modeling and the text region semantic modeling. By capturing the context information in the segment relationship modeling module, discriminative cross-modal embedding knowledge is learned, and then the overall embedding is learned.
[0011] Step 3: Construct a multi-scale feature enhancement module and add it to the cross-modal pedestrian re-identification backbone network in Step 2 to enhance the perception of the target object by the cross-modal pedestrian re-identification backbone network and obtain richer pedestrian image features;
[0012] Step 4: Determine the loss function to complete the final training of the cross-modal pedestrian re-identification backbone network;
[0013] Step 5: Use one-modal pedestrian images as the query set and the other-modal pedestrian images as the retrieval set, and match the pedestrian images in the query set with the pedestrian images in the retrieval set to obtain the cross-modal pedestrian re-identification result.
[0014] A further improvement of the present invention is that: in the said Step 2, the cross-modal pedestrian re-identification backbone network further includes a visible light texture encoder, a joint relationship encoder, and an infrared texture encoder. The visible light image and the infrared light image pass through the constructed cross-modal pedestrian re-identification backbone network. The paths of the visible light image and the infrared light image are respectively: the visible light path, the infrared path, and the visible light-infrared shared path:
[0015] For the visible light path, the input is the visible light image. The process of the visible light image passing through the cross-modal pedestrian re-identification backbone network is as follows: the visible light image enters the color encoder and passes through the visual area image modeling module to obtain color features, which enter the visible light joint relationship encoder together with the texture features obtained through the visible light texture encoder for full fusion; For the infrared path, the input is the image features after supplementing the missing color information by the text pre-training model in Step 1. The process of passing through the cross-modal pedestrian re-identification backbone network is as follows: it enters the text encoder
[0016] and passes through the text area semantic modeling module to obtain text features, which enter the infrared joint relationship encoder together with the texture features obtained through the infrared texture encoder for full fusion; For the visible light-infrared shared path, the process is as follows: the texture features obtained from the visible light path and the infrared path enter the feature enhancement network with shared parameters to improve the regional perception ability of the cross-modal pedestrian re-identification backbone network for visible light images and infrared images, and obtain more useful spectral information through the channel spectrum change module
[0017] and finally be constrained by the loss function. to improve the regional perception ability of the cross-modal pedestrian re-identification backbone network for visible light images and infrared images, and obtain more useful spectral information through the channel spectrum change module and finally be constrained by the loss function.
[0018] A further improvement of the present invention lies in that: the global-local cross-attention mechanism in step 2 learns discriminative cross-modal embedding knowledge, and then learns the overall embedding, which specifically includes the following steps:
[0019] Step 2.1.1: By calculating the product of the attention matrices from the lower layer to the upper layer of the segment relationship modeling module, calculate the cumulative attention score of this specified block of the region from the lower layer to the upper layer of the segment relationship modeling module, and obtain the attention map. The calculation process is described as follows:
[0020]
[0021] Among them, represents the re-normalization of the attention weights, S represents the weighted attention matrix of each layer, represents the matrix multiplication operation, 1-i represents from the first layer to the i-th layer, and thus the information when propagating from the input layer to the higher layer;
[0022] Step 2.1.2: Mine the high-response region by aggregating the attention map, is expressed as the cumulative weighting of the class embedding, and select the top R query vectors from the query matrix Q i to reconstruct the new query matrix Q l to represent the most concerned local embedding information;
[0023] Step 2.1.3: From the constructed new query matrix Q l and the global set of the entire key-value pair, calculate the cross-attention to obtain the attention weight matrix. The specific calculation process is as follows:
[0024]
[0025] Among them, d is the size dimension of the input feature, softmax is the activation function operation, K gT is the new query vector matching the query vector, V is the value vector, and Y att represents the final cross-attention output weight.
[0026] A further improvement of the present invention lies in that: in step 2, the visual region image modeling module includes three parts: region mapping, region feature semantics, and sample segment interaction, which specifically includes the following steps:
[0027] Step 2.2.1: The input coal mine infrared light image and coal mine visible light image dynamically adjust the dimension size of the embedding vector through the region mapping part, and add a fully connected layer to map the mapped region features to d-dimensional local features as visual segments, and denote them as the initial visible light region features where n r is the number of regional features;
[0028] Step 2.2.2, the obtained initial regional features Construct a semantic fully-connected relationship graph network feature representing the visual region, obtain the local spatial information of the target in the coal mine visible light image, and learn the context semantics of the initial regional features through relationship interaction. Among them, the graph nodes represent the visible light initial regional feature R, the edges represent their corresponding semantic relationships, and the global-local cross-attention mechanism is used to capture the semantic relationships of the visible light initial regional feature R and learn the mutually related enhanced local features, and the semantic relationships are hidden into the attention weights to obtain the relationship-enhanced regional features
[0029] Step 2.2.3, obtain the global visual embedding feature Y by aggregating the visible light initial regional feature R and the relationship-enhanced regional feature v , and its calculation process can be described as:
[0030] Y v = λ·MaxPool(R)+(1 - λ)·AvgPool(R v )
[0031] where, the visible light initial regional feature R uses max pooling MaxPool, the enhanced regional feature uses average pooling AvgPool operation, and λ is the proportionality coefficient.
[0032] A further improvement of the present invention lies in: in step 2, the text region semantic modeling includes three parts: region mapping, lexical feature semantics, and sample segment interaction, and specifically includes the following steps:
[0033] Step 2.3.1, input the coal mine infrared light image and the coal mine visible light image to dynamically adjust the dimension size of the embedding vector through the region mapping part, and add a fully-connected layer to map the regional feature to a d-dimensional local feature as the text segment, and denote it as the infrared light initial regional feature where n t is the number of regional features;
[0034] Step 2.3.2, the obtained infrared light initial regional features Construct a word connection relationship graph network representing the text region, obtain the local position information of the target in the word, and utilize the lexical dependency information and position information through relationship interaction. Among them, the graph nodes represent the infrared light initial regional feature T, the edges represent their corresponding semantic relationships, and the global-local cross-attention mechanism is used to capture their semantic relationships and learn the mutually related enhanced local features, and the semantic relationships are hidden into the attention weights to obtain the relationship-enhanced lexical features
[0035] Step 2.3.3: Obtain the global visual embedding feature U by aggregating the initial infrared light region feature T and the relationship-enhanced vocabulary feature v , and its calculation process can be described as follows:
[0036] U v = λ·MaxPool(T)+(1 - λ)·AvgPool(T v )
[0037] where the initial infrared light region feature T uses max pooling MaxPool, the relationship-enhanced vocabulary feature uses average pooling AvgPool operation, and λ is the proportionality coefficient.
[0038] A further improvement of the present invention lies in: in the step 2, a channel spectrum transformation module is added to the global-local cross-attention mechanism, and discrete cosine transform is performed through the channel spectrum transformation module. The specific transformation process includes the following steps:
[0039] Step 2.4.1: Divide the input X i along the channel dimension into multiple parts and represent them as [X 0 , X 0 , …, X n-1 , where X i ∈R C′×H×W , i ∈ {0, 1, ..., k - 1}, and assign corresponding two-dimensional discrete cosine transform frequency components to each part to obtain the result Y of the discrete cosine transform;
[0040] Step 2.4.2: Take the obtained result of the discrete cosine transform as the compression result and represent it as:
[0041]
[0042] where i ∈ {0, 1, ..., k - 1}, [u i , v i corresponds to the corresponding 2D index of the frequency component, is the calculated discrete cosine value, is the input value of the input X i at the specified height h and width w, and Y freq is the compressed C'-dimensional vector, and the global compressed vector is obtained through concatenation and represented as:
[0043]
[0044] where cat means concatenating multiple tensors together;
[0045] Step 2.4.3. Obtain the weights of each channel through the fully connected layer and the output after passing through the spectral channel, which is specifically expressed as:
[0046] Y att = sigmoid(fc(Y freq ))
[0047] where sigmoid represents a smooth and monotonic activation function that maps any real value to the interval (0, 1).
[0048] A further improvement of the present invention lies in that: the multi-scale feature enhancement module in Step 3 adopts a convolutional generation structure with three different branches, each branch structure has different convolutional configurations, extracts multiple distinguishable semantic information, enhances the feature representation of the target, and performs feature enhancement through the multi-scale feature enhancement module. The specific process is as follows:
[0049] Step 3.1. The input feature f i enters three different branches respectively and obtains three different outputs F1, F2, F3. Its calculation process can be described as:
[0050]
[0051] where represent convolutional operations with convolutional kernel sizes of 3×3, 1×1, 1×3, and 3×1 respectively, represents a dilated convolutional operation with a convolutional kernel size of 3×3 and a dilation rate of 5;
[0052] Step 3.2. After the input feature passes through three different branches, it is merged and concatenated with the residual connection after passing through a 1×1 convolution to obtain the final output f out , and its calculation process can be described as:
[0053]
[0054] where F cat represents the concatenation operation on the feature map, represents a standard convolutional operation with a convolutional kernel size of 1×1.
[0055] A further improvement of the present invention lies in that: in Step 4, the orthogonal projection loss function, cross-entropy loss function, triplet loss function, center-guided pair mining loss function, and KL divergence loss function are used to jointly constrain the model training process, the encoder and the features after the encoder are optimized by the KL divergence loss function, and through the encoder and The features after the encoder are optimized by the triplet loss function, and the features after the visible-infrared shared path are jointly optimized by the center-guided pair mining loss function, the cross-entropy loss function, and the orthogonal projection loss function.
[0056] A further improvement of the present invention lies in that: the calculation process of the orthogonal projection loss function is as follows:
[0057] First, define the calculation formula of the cosine similarity operator for two vectors, which is specifically shown as follows:
[0058]
[0059] where, ‖g‖2 represents the calculation process of the norm l2 operator norm;
[0060] Secondly, the clustering calculation process for samples of the same category and the orthogonality calculation process for samples of different categories are defined as follows:
[0061]
[0062] where, <x i ,y i > is the input-output pair, and calculate the cosine similarity operator of two vectors, represents the output of the middle layer of the network, and B represents the size of the mini-batch;
[0063] Finally, the total orthogonal projection loss function is expressed as:
[0064] L ort =(1 - d)+|h|
[0065] where, |·| is the absolute value operator.
[0066] The cross-entropy loss function is used for identity recognition in the pedestrian retrieval process. Regarding the cross-modal pedestrian re-identification task as an image classification task, its calculation process can be expressed as:
[0067]
[0068] where, n represents the number of training samples, x i represents the input of the pedestrian image, y i represents its corresponding identity label, and p(y i |x i ) represents the predicted probability that the input image and its corresponding identity label are recognized as after being classified by the Softmax function;
[0069] The triplet loss function is used to optimize the triplet relationship between the same identity and different identities, ensuring that the feature distance of the same pedestrian image is less than the feature distance between different pedestrians. The cross-modal pedestrian re-identification task is regarded as an image retrieval task, and its calculation process can be expressed as:
[0070]
[0071] Among them, and represent the anchor sample, the positive sample, and the negative sample respectively. and belong to the same pedestrian ID. and respectively use the IDs of different pedestrians. F(·) represents the feature extraction function. represents the process of solving the Euclidean distance between the positive and negative samples. The most difficult positive sample is the sample that belongs to the same category as the anchor sample and has the farthest distance, and the most difficult negative sample is the sample that belongs to a different category from the anchor sample and has the closest distance. P represents the number of pedestrians included in a single batch, K represents the number of corresponding pictures for each pedestrian, and α is the margin parameter.
[0072] The center-guided pair mining loss function is used to constrain the generated embedding information. The generated embedding is constrained by the following three attribute constraints, pulling the distance between the generated embedding and the original embedding, narrowing the distance between the generated embedding of the visible light image and the original embedding of the infrared light image, and narrowing the distance between the generated embedding of the original image and the original embedding of the visible light image, ensuring that the intra-class distance is less than the inter-class distance. Among them, α is the balance parameter, and its calculation process is expressed as:
[0073]
[0074] Similarly, for the control of the class center of the embedding generated from the infrared light image, its calculation process can be expressed as:
[0075]
[0076] The final center-guided pair mining loss function can be expressed as:
[0077]
[0078] The calculation process of the KL divergence loss function is:
[0079]
[0080] Among them, p n,m is the pairing probability of {f n , f m}, q n,m is its true matching probability, and the constant ε is set to 1e-8 .
[0081] The beneficial effects of the present invention are:
[0082] Aiming at the scarcity of pedestrian samples in existing coal mine scenarios, a fragment relationship modeling architecture is proposed, focusing on the fragment-level interaction between samples, effectively capturing the contextual information of pedestrian fragments, and enhancing the feature representation of the network;
[0083] The present invention proposes a channel spectrum transformation process, which compresses the channel and utilizes multiple frequency components of discrete cosine transform to mine potential useful frequency information that is not noticed by general channel attention;
[0084] The present invention proposes a multi-scale feature enhancement module, which makes full use of local information and global context information, fully combines high-resolution low-level features and low-resolution high-level features, effectively enhances the network's perception of the target object, and obtains richer pedestrian image features.
[0085] In summary, the present invention effectively enhances the interaction between pedestrian cross-modal features and mines richer pedestrian image features through relational modeling and spectral transformation, effectively improving the prediction and generalization capabilities of the model in cross-modal pedestrian image retrieval scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is a flow chart of the cross-modal pedestrian re-identification method of the present invention.
[0087] Figure 2 Schematic diagram of the cross-modal person re-identification backbone network of the present invention.
[0088] Figure 3 It is a structural diagram of the cross-modal pedestrian re-identification backbone network of the present invention.
[0089] Figure 4 It is a schematic diagram of the visual area image modeling module of the present invention.
[0090] Figure 5 It is a schematic diagram of the channel spectrum conversion process of the present invention.
[0091] Figure 6 Schematic diagram of the multi-scale feature enhancement module of the present invention. DETAILED DESCRIPTION
[0092] The following will disclose the embodiments of the present invention with drawings. For the purpose of clear description, many practical details will be described together in the following description. However, it should be understood that these practical details should not be used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0093] In view of the complex environment in the coal mine operation area, which is easily affected by factors such as soot and insufficient light, the present invention effectively extracts pedestrian image features and reduces the differences between modalities by adding a fragment relationship modeling structure and a channel spectrum transformation process. Specifically, a cross-modal pedestrian re-identification method based on a coal mine scenario is proposed, as Figure 1-2 shown, and the method specifically includes the following steps:
[0094] Step 1: Obtain the identity representing the pedestrian image and the identity label corresponding to the pedestrian image from the coal mine infrared light image and the coal mine visible light image. Each pedestrian image of different modalities includes a visible light image and an infrared light image, and preprocess the visible light image and the infrared light image. Pass the visible light image and the infrared light image through a text pre-training model for preprocessing to obtain the visible light image features and infrared light image features after the text pre-training model supplements the missing color information;
[0095] Step 2: Construct a cross-modal pedestrian re-identification backbone network based on relationship modeling and spectrum transformation. The cross-modal pedestrian re-identification backbone network mainly includes: a color encoder, a text encoder, a fragment relationship modeling module, a channel spectrum transformation module, a visible light texture encoder, a joint relationship encoder, and an infrared texture encoder; wherein, the fragment relationship modeling module includes a visual region image modeling module and a text region semantic modeling module, and a global-local cross-attention mechanism is added to each part of the visual region image modeling and the text region semantic modeling. By capturing the context information in the fragment relationship modeling module, discriminative cross-modal embedding knowledge is learned, and then the overall embedding is learned.
[0096] The overall network framework is as Figure 3 shown. The cross-modal pedestrian re-identification backbone network also includes a visible light texture encoder, a joint relationship encoder, and an infrared texture encoder. The visible light image and the infrared light image pass through the constructed cross-modal pedestrian re-identification backbone network. The paths of the visible light image and the infrared light image are respectively: the visible light path, the infrared path, and the visible light-infrared shared path:
[0097] For the visible light path, the input is the visible light image. The process of the visible light image passing through the cross-modal pedestrian re-identification backbone network is as follows: The visible light image enters the color encoder and obtains color features through the visual region image modeling module and jointly enters the visible light joint relationship encoder with the texture features obtained through the visible light texture encoder for full fusion; for full fusion;
[0098] For the infrared path, the input is the image features after supplementing the missing color information by the text pre-training model in step 1. The process passing through the cross-modal person re-identification backbone network is as follows: Enter the text encoder and pass through the text region semantic modeling module to obtain text features, which, together with the texture features obtained by the infrared texture encoder enter the infrared joint relationship encoder for full fusion;
[0099] For the visible-infrared shared path, the process is as follows: The texture features obtained from the visible light path and the infrared path enter the feature enhancement network with shared parameters to improve the regional perception ability of the cross-modal person re-identification backbone network for visible light images and infrared images, and obtain more useful spectral information through the channel spectrum change module Finally, it is constrained by the loss function.
[0100] The global-local cross-attention mechanism obtains richer image relationships and semantic relationships by strengthening the interaction between the global image and the local significant regions, learns discriminative cross-modal embedding knowledge, and then learns the overall embedding. It specifically includes the following steps:
[0101] Step 2.1.1: Calculate the cumulative attention score of this specified block of the region from the low layer to the high layer of the segment relationship modeling module by calculating the product of the attention matrices from the low layer to the high layer of the segment relationship modeling module, and obtain the attention map. The calculation process is described as:
[0102]
[0103] Among them, represents the re-normalization of the attention weights, S represents the weighted attention matrix of each layer, represents the matrix multiplication operation, 1 - i represents from the first layer to the i-th layer, and thus the information when propagating from the input layer to higher layers;
[0104] Step 2.1.2: Mine the high-response regions by aggregating the attention map, represented as the cumulative weighting of the class embeddings, and select the top R query vectors from the query matrix Q i to reconstruct a new query matrix Q l to represent the most concerned local embedding information;
[0105] Step 2.1.3: From the constructed new query matrix Q lFor the global set of all key-value pairs, cross-attention is calculated to obtain the attention weight matrix. The specific calculation process is as follows:
[0106]
[0107] where d is the dimensionality of the input features, softmax is the activation function operation, K gT is the new query vector that matches the query vector, V is the value vector, and Y att represents the final cross-attention output weight.
[0108] As Figure 4 shown, the visual region image modeling module includes three parts: region mapping, region feature semantics, and sample segment interaction. The region feature semantics part learns the attention weights of different image relationships in the region features through the image relationship graph between visual regions, and learns the context information of the region features and extracts image relationships through the global-local cross-attention mechanism. The sample segment interaction part captures the relationship features between samples through the max pooling operation and the average pooling operation, and then learns the relationship-enhanced pedestrian image features. The specific steps are as follows:
[0109] Step 2.2.1: The input coal mine infrared light image and coal mine visible light image dynamically adjust the dimension size of the embedding vector through the region mapping part, and add a fully connected layer. The mapped region features are mapped to d-dimensional local features as visual segments, and are denoted as the initial visible light region features where n r is the number of region features;
[0110] Step 2.2.2: The obtained initial region features construct the semantic fully connected relationship graph network features representing the visual regions, obtain the local spatial information of the targets in the coal mine visible light image, and learn the context semantics of the initial region features through relationship interaction. Among them, the graph nodes represent the initial visible light region features R, the edges represent their corresponding semantic relationships, and the global-local cross-attention mechanism is used to capture the semantic relationships of the initial visible light region features R and learn the mutually relationship-enhanced local features, and the semantic relationships are hidden into the attention weights to obtain the relationship-enhanced region features
[0111] Step 2.2.3: Aggregate the initial visible light region features R and the relationship-enhanced region features to obtain the global visual embedding feature Y v , and its calculation process can be described as:
[0112] Y v = λ·MaxPool(R)+(1 - λ)·AvgPool(R v )
[0113] Among them, for the initial visible light region feature R, max pooling MaxPool is used, and for the enhanced region feature, average pooling AvgPool operation is used, where λ is a proportionality coefficient.
[0114] The text region semantic modeling includes three parts: region mapping, lexical feature semantics, and sample segment interaction. In the lexical feature semantics part, a bidirectional language representation result is generated through the text relationship graph within the text region, and important semantic information is aggregated through the global-local cross-attention mechanism to filter out irrelevant information and perform text relationship extraction; in the sample segment interaction part, the relationship features between samples are captured through max pooling operation and average pooling operation to learn the relationship-enhanced human text features. The specific steps are as follows:
[0115] Step 2.3.1: Input the coal mine infrared light image and the coal mine visible light image to dynamically adjust the dimension size of the embedding vector through the region mapping part, and add a fully connected layer to map the region feature to a d-dimensional local feature as a text segment, denoted as the initial infrared light region feature where n t is the number of region features;
[0116] Step 2.3.2: The obtained initial infrared light region feature Construct a word connection relationship graph network representing the text region, obtain the local position information of the target in the word, and utilize the lexical dependency information and position information through relationship interaction. Among them, the graph node represents the initial infrared light region feature T, the edge represents its corresponding semantic relationship, and the global-local cross-attention mechanism is used to capture its semantic relationship and learn the mutually relationship-enhanced local feature, and the semantic relationship is hidden into the attention weight to obtain the relationship-enhanced lexical feature
[0117] Step 2.3.3: Aggregate the initial infrared light region feature T and the relationship-enhanced lexical feature to obtain the global visual embedding feature U v , and its calculation process can be described as:
[0118] U v = λ·MaxPool(T)+(1 - λ)·AvgPool(T v )
[0119] Among them, for the initial infrared light region feature T, max pooling MaxPool is used, and for the relationship-enhanced lexical feature, average pooling AvgPool operation is used, where λ is a proportionality coefficient.
[0120] The present invention adds a channel spectrum transformation module, namely the discrete cosine transform process, to the global-local cross attention mechanism, and designs a channel spectrum transformation process that can be applied to the field of cross-modal person re-identification. While compressing the channel, it utilizes multiple frequency components of the discrete cosine transform to mine potential useful frequency information that is not noticed by general channel attention. Figure 5 As shown, the input feature vector uses multiple frequency components based on discrete cosine transform, and provides frequency component selection criteria for selection, and finally obtains more information including multiple frequency components.
[0121] The specific transformation process includes the following steps:
[0122] Step 2.4.1. Input X i It is divided into multiple parts along the channel dimension and expressed as [X 0 ,X 0 ,…,X n-1 ], where X i ∈R C′×H×W , i∈{0, 1, …, k-1}, And assign the corresponding two-dimensional discrete cosine change frequency component to each part, and get the discrete cosine transform result Y. Figure 2 As shown, enter X i is the output of the texture encoder;
[0123] Step 2.4.2: The obtained discrete cosine transform result is used as the compression result and expressed as:
[0124]
[0125] where i∈{0, 1, …, k-1}, [u i ,v i ] the corresponding 2D index of the corresponding frequency component, is the calculated discrete cosine value, For input X i When the input value of height h and width w is specified, Y freq That is the compressed C'-dimensional vector, and the global compressed vector is obtained by cascading, which is expressed as:
[0126]
[0127] Among them, cat means concatenating multiple tensors together;
[0128] Step 2.4.3, after obtaining the weight of each channel through the fully connected layer and passing through the spectral channel, the output is specifically expressed as:
[0129] Y att =sigmoid(fc(Y freq))
[0130] Among them, sigmoid represents a smooth and monotonic activation function that maps any real value to the interval (0, 1).
[0131] Step 3: Construct a multi-scale feature enhancement module and add it to the cross-modal pedestrian re-identification backbone network in Step 2 to enhance the perception of the target object by the cross-modal pedestrian re-identification backbone network and obtain richer pedestrian image features. As Figure 6 shown, the multi-scale feature enhancement module adopts a convolutional generation structure with three different branches. Each branch structure has different convolutional configurations, extracts multiple distinguishable semantic information, enhances the feature representation of the target, and performs feature enhancement through the multi-scale feature enhancement module. The specific process is as follows:
[0132] Step 3.1: The input feature f i enters three different branches respectively and obtains three different outputs F1, F2, F3. Its calculation process can be described as:
[0133]
[0134] Among them, respectively represent convolutional operations with a convolutional kernel size of 3×3, 1×1, 1×3, and 3×1, represents a dilated convolution operation with a convolutional kernel size of 3×3 and a dilation rate of 5; according to the appendix Figure 2 shown, the feature f i is the output obtained through the texture encoder and is used as the input of the multi-scale feature enhancement module here.
[0135] Step 3.2: After the input feature passes through three different branches, it is merged and concatenated with the residual connection after passing through a 1×1 convolution to obtain the final output f out , and its calculation process can be described as:
[0136]
[0137] Among them, F cat represents the concatenation operation on the feature map, represents a standard convolution operation with a convolutional kernel size of 1×1.
[0138] Step 4: Determine the loss function, training environment, and training details, complete the final training of the cross-modal pedestrian re-identification backbone network, use the pedestrian images of one modality as the query set, the pedestrian images of the other modality as the retrieval set, match the pedestrian images in the query set with the pedestrian images in the retrieval set, and obtain the cross-modal pedestrian re-identification result.
[0139] The present invention uses an orthogonal projection loss function, a cross-entropy loss function, a triplet loss function, a center-guided pair mining loss function, and a KL divergence loss function to jointly constrain the model training process. The encoder and The features after the encoder are optimized by the KL divergence loss function. Through The encoder and The features after the encoder are optimized by the triplet loss function, and the features after the visible-infrared shared path are jointly optimized by the center-guided pair mining loss function, the cross-entropy loss function, and the orthogonal projection loss function.
[0140] The present invention uses an orthogonal projection loss function. By implementing an orthogonality constraint in the intermediate feature space and using a projection method, it explores the possibility of implementing, in a small batch, measures aimed at enhancing the intrinsic discriminative features between samples, maximally separating the different features generated by different classes, supplementing the angular discriminability of the output space, and further enhancing the discriminative and generalization capabilities of the model in learning feature representations during the training process.
[0141] First, the calculation formula for the cosine similarity operator of two vectors is defined as follows: as shown in the following formula:
[0142]
[0143] where, ‖g‖2 represents the calculation process of the norm l2 operator norm;
[0144] Second, the clustering calculation process for samples of the same class and the orthogonality calculation process for samples of different classes are defined as follows:
[0145]
[0146] where, <x i , y i > is the input-output pair, and the cosine similarity operator of two vectors is calculated. represents the output of the intermediate layer of the network, and B represents the size of the small batch;
[0147] Finally, the total orthogonal projection loss function is expressed as:
[0148] L ort =(1 - d)+|h|
[0149] where, |·| is the absolute value operator.
[0150] The orthogonal projection loss function utilizes efficient vectorized implementation, direct gradient calculation and propagation, and orthogonalization in the feature space to optimize different objectives of between-class separation and within-class clustering. By clustering different within-class samples simultaneously, it enforces all classes to be orthogonal to each other. Maximizing the difference may lead to negative correlations between classes, ensuring that the embeddings generated from different branches can capture different information features.
[0151] The cross-entropy loss function prevents the model from overfitting by enhancing the discriminative ability of identity information. This function is used for identity recognition in the pedestrian retrieval process during training. Treating the cross-modal pedestrian re-identification task as an image classification task, its calculation process can be expressed as:
[0152]
[0153] where n represents the number of training samples, x i represents the input of the pedestrian image, y i represents its corresponding identity label, and p(y i |x i ) represents the predicted probability that the input image and its corresponding identity label are recognized as after being classified by the Softmax function;
[0154] The triplet loss function is used to optimize the triplet relationship between the same identity and different identities, ensuring that the feature distance of the same pedestrian image is less than that between different pedestrians. Treating the cross-modal pedestrian re-identification task as an image retrieval task, its calculation process can be expressed as:
[0155]
[0156] where, and represent the anchor sample, positive sample, and negative sample respectively, and belong to the ID of the same pedestrian, and use the IDs of different pedestrians respectively. F(·) represents the feature extraction function, represents the process of solving the Euclidean distance between the positive and negative samples. The most difficult positive sample is the sample that belongs to the same category as the anchor sample and is the farthest away, and the most difficult negative sample is the sample that belongs to a different category from the anchor sample and is the closest. P represents the number of pedestrians included in a single batch, K represents the number of corresponding pictures for each pedestrian, and α is the margin parameter, set to 0.3;
[0157] The central guidance pair mining loss function is used to constrain the generated embedding information. The generated embedding is constrained by the following three property constraints to increase the distance between the generated embedding and the original embedding, and decrease the distances between the generated embedding of the visible light image and the original embedding of the infrared light image, and between the generated embedding of the original image and the original embedding of the visible light image, ensuring that the intra-class distance is less than the inter-class distance. Here, α is a balance parameter set to 0.2. Its calculation process is expressed as:
[0158]
[0159] Similarly, for the control of the class center of the embedding generated from the infrared light image, its calculation process can be expressed as:
[0160]
[0161] The final central guidance pair mining loss function can be expressed as:
[0162]
[0163] The calculation process of the KL divergence loss function is:
[0164]
[0165] Among them, p n,m is the pairing probability of {f n , f m}, q n,m is its true matching probability, and the constant ε is set to 1e -8 .
[0166] The present invention jointly uses an orthogonal projection loss function, a cross-entropy loss function, a triplet loss function, a central guidance pair mining loss function, and a KL divergence loss. Among them, the encoder and the features after the encoder are optimized by the KL divergence loss function, and the features after the encoder and the encoder are optimized by the triplet loss function. The features after the visible light-infrared shared path are jointly optimized by the central guidance pair mining loss function, the cross-entropy loss function, and the orthogonal projection loss function.
[0167] During the training phase, the size of the input images is uniformly adjusted to 384×144. Each training mini-batch consists of 4 identities, and each identity contains 4 visible light images and 4 infrared images. Random cropping, random horizontal flipping, and random erasing are performed for data augmentation. The number of training iterations is 80 epochs. The sgd optimizer is used for optimization, and the learning rate is set to 0.1 in the first round. At the 20th epoch and 50th epoch, the learning rate is decayed to 0.01 and 0.001 in sequence.
[0168] When training on the RegDB dataset, since there are relatively few pedestrian images, the first 3 stages of the ResNet-50 structure are used for training. A fragment relationship modeling structure and a channel spectrum transformation process are added accordingly, and a feature enhancement module is added in the second stage.
[0169] The technical effects of the present invention will be described in detail below in combination with performance tests. The present invention aims to improve the prediction ability and generalization ability of the model in the cross-modal pedestrian image retrieval scenario. To prove the performance of this model in the coal mine scenario, the present invention uses a cross-modal pedestrian re-identification dataset for sufficient testing and compares it with existing algorithms in recent years. The results are shown in the following table:
[0170] Table 1 Comparison of results with existing methods on the SYSU-MM01 dataset (%)
[0171]
[0172]
[0173] Under All-search in the dataset, the R-1, R-10, R-20, and mAP of the method in this paper reached 84.8%, 99.1%, 99.9%, and 81.5%, 1.5% respectively. Under Indoor-search, the R-1, R-10, R-20, and mAP reached 92.5%, 99.5%, 100.0%, and 92.6% respectively, fully confirming the effectiveness of the method.
[0174] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.
Claims
1. A cross-modal pedestrian re-identification method based on coal mine scenes, characterized by: The coal mine scene cross-modal pedestrian re-identification method specifically includes the following steps: Step 1: Obtain the identity of the pedestrian image and the identity label corresponding to the pedestrian image from the coal mine infrared image and the coal mine visible light image. Each pedestrian image of different modalities contains a visible light image and an infrared light image. The visible light image and the infrared light image are preprocessed. The visible light image and the infrared light image are preprocessed by a text pre-training model to obtain the visible light image features and the infrared light image features after the text pre-training model supplements the missing color information; Step 2: construct a cross-modal pedestrian re-identification backbone network based on relational modeling and spectrum transformation, wherein the cross-modal pedestrian re-identification backbone network mainly includes: a color encoder, a text encoder, a fragment relational modeling module, a channel spectrum transformation module, a visible light texture encoder, a joint relational encoder, and an infrared texture encoder; wherein the fragment relational modeling module includes a visual area image modeling module and a text area semantic modeling module, and a global-local cross attention mechanism is added to each part of the visual area image modeling and the text area semantic modeling, and by capturing the contextual information in the fragment relational modeling module, discriminative cross-modal embedding knowledge is learned, and then the overall embedding is learned; Step 3: construct a multi-scale feature enhancement module and add it to the cross-modal pedestrian re-identification backbone network in step 2 to enhance the perception of the cross-modal pedestrian re-identification backbone network to the target object and obtain richer pedestrian image features; Step 4: Determine the loss function and complete the final training of the cross-modal person re-identification backbone network; Step 5: Use one of the modal pedestrian images as the query set and the other modal pedestrian images as the retrieval set, match the pedestrian images in the query set with the pedestrian images in the retrieval set, and obtain the cross-modal pedestrian re-identification results.
2. According to the coal mine scene-based cross-modal pedestrian re-identification method of claim 1, it is characterized by: In step 2, the cross-modal pedestrian re-identification backbone network also includes a visible light texture encoder, a joint relationship encoder, and an infrared texture encoder. The visible light image and the infrared light image are constructed through a cross-modal pedestrian re-identification backbone network, and the paths of the visible light image and the infrared light image are respectively: a visible light path, an infrared path, and a visible light-infrared shared path: For the visible light path, the input is a visible light image. The process of the visible light image passing through the cross-modal pedestrian re-identification backbone network is as follows: the visible light image enters the color encoder And through the visual area image modeling module Obtain color features and use the visible light texture encoder The acquired texture features are jointly fed into the visible light joint relation encoder to integrate; For the infrared path, the input is the image features after the missing color information is supplemented by the text pre-trained model in step 1. The process of passing through the cross-modal pedestrian re-identification backbone network is as follows: Enter the text encoder And through the text region semantic modeling module Get text features and use infrared texture encoder The acquired texture features are jointly fed into the infrared joint relation encoder to integrate; For the visible light-infrared shared path, the process is as follows: the texture features obtained by the visible light path and the infrared path are fed into the parameter-sharing feature enhancement network Improve the regional perception ability of the cross-modal pedestrian re-identification backbone network for visible light images and infrared images, and use the channel spectrum change module Obtain more useful spectrum information and finally constrain it through the loss function.
3. According to the coal mine scene-based cross-modal pedestrian re-identification method of claim 1, it is characterized by: The global-local cross attention mechanism in step 2 learns discriminative cross-modal embedding knowledge, and then learns the overall embedding, which specifically includes the following steps: Step 2.1.1, by calculating the product of the attention matrix from the low layer of the fragment relationship modeling module to the high layer of the fragment relationship modeling module, the cumulative attention score of the specified block in the area from the low layer of the fragment relationship modeling module to the high layer of the fragment relationship modeling module is calculated to obtain the attention map. The calculation process is described as: in, represents the renormalization of the attention weights, S represents the weighted attention matrix of each layer, Represents matrix multiplication operation, 1-i represents from the 1st layer to the i-th layer, thus, the information propagated from the input layer to the higher layer; Step 2.1.2: Mining high response areas by aggregating attention maps. It is represented as the cumulative weighted class embedding and is obtained from the query matrix Q i Select the query vectors ranked in the top R to reconstruct the new query matrix Q l , used to represent the most concerned local embedded information; Step 2.1.3: The new query matrix Q constructed by l And the global set of the entire key-value pair, calculate the cross attention to get the attention weight matrix. The specific calculation process is: Among them, d is the dimension of the input feature, softmax is the activation function operation, K gT is the new query vector that matches the query vector, V is the value vector, and Y att Represented as the final cross-attention output weight.
4. The cross-modal person re-identification method based on coal mine scene according to claim 1 is characterized by: In step 2, the visual region image modeling module includes three parts: region mapping, region feature semantics, and sample segment interaction, and specifically includes the following steps: Step 2.2.1: The input coal mine infrared image and coal mine visible light image are embedded in the region mapping part to dynamically adjust the dimension size of the vector, and add a fully connected layer to map the mapping region feature to the d-dimensional local feature as a visual fragment, which is recorded as the visible light initial region feature. Where n r is the number of regional features; Step 2.2.2: Obtained initial region features A semantic fully connected relational graph network feature representing the visual area is constructed to obtain the local spatial information of the target in the visible light image of the coal mine, and the contextual semantics of the initial regional features are learned through relational interaction. The graph nodes represent the visible light initial regional features R, and the edges represent their corresponding semantic relations. The global-local cross attention mechanism is used to capture the semantic relations of the visible light initial regional features R and learn the local features enhanced by mutual relations. The semantic relations are implied in the attention weights to obtain the relation-enhanced regional features. Step 2.2.3: Obtain the global visual embedding feature Y by aggregating the initial visible light region feature R and the relationship enhanced region feature v , and its calculation process can be described as: Y v =λ·MaxPool(R)+(1-λ)·AvgPool(R v ) Among them, the initial region feature R of visible light uses the maximum pooling MaxPool, the enhanced region feature uses the average pooling AvgPool operation, and λ is the proportional coefficient.
5. The cross-modal person re-identification method based on coal mine scene according to claim 1 is characterized by: In step 2, the text region semantic modeling includes three parts: region mapping, vocabulary feature semantics, and sample segment interaction, and specifically includes the following steps: Step 2.3.
1. Input the infrared image and visible light image of the coal mine, dynamically adjust the embedding vector dimension through the region mapping part, and add a fully connected layer to map the regional features to d-dimensional local features as text fragments, and record them as infrared initial regional features. Where n t is the number of regional features; Step 2.3.2: Obtained infrared light initial area features A word connection relationship graph network representing the text area is constructed to obtain the local position information of the target in the word, and the word dependency information and position information are utilized through relationship interaction. The graph nodes represent the initial regional features T of infrared light, and the edges represent the corresponding semantic relations. The global-local cross attention mechanism is used to capture the semantic relations and learn the local features enhanced by mutual relations. The semantic relations are implied in the attention weights to obtain the relationship-enhanced vocabulary features. Step 2.3.3: Obtain the global visual embedding feature U by aggregating the infrared initial region feature T and the relation-enhanced vocabulary feature v , and its calculation process can be described as: U v =λ·MaxPool(T)+(1-λ)·AvgPool(T v ) Among them, the infrared light initial area feature T uses the maximum pooling MaxPool, the relationship enhancement vocabulary feature uses the average pooling AvgPool operation, and λ is the proportional coefficient.
6. The cross-modal person re-identification method based on coal mine scene according to claim 1 is characterized by: In step 2, a channel spectrum transformation module is added to the global-local cross attention mechanism, and discrete cosine transform is performed through the channel spectrum transformation module. The specific transformation process includes the following steps: Step 2.4.
1. Input X i It is divided into multiple parts along the channel dimension and expressed as [X 0 ,X 0 ,…,X n-1 ], where X i ∈R C '×H×W , i∈{0, 1, ..., k-1}, And assign the corresponding two-dimensional discrete cosine transform frequency component to each part to obtain the discrete cosine transform result Y; Step 2.4.2: The obtained discrete cosine transform result is used as the compression result and expressed as: where i∈{0, 1, ..., k-1}, [u i ,v i ] the corresponding 2D index of the corresponding frequency component, is the calculated discrete cosine value, For input X i When the input value of height h and width w is specified, Y freq That is the compressed C'-dimensional vector, and the global compressed vector is obtained by cascading, which is expressed as: Among them, cat means concatenating multiple tensors together; Step 2.4.3, after obtaining the weight of each channel through the fully connected layer and passing through the spectral channel, the output is specifically expressed as: Y att =sigmoid(fc(Y freq )) Among them, sigmoid represents a smooth and monotonic activation function that maps any real value to the interval (0,1).
7. The cross-modal person re-identification method based on coal mine scene according to claim 1 is characterized by: The multi-scale feature enhancement module in step 3 adopts a convolutional generation structure with three different branches. Each branch structure has a different convolution configuration, extracts multiple distinguishable semantic information, enhances the feature representation of the target, and performs feature enhancement through the multi-scale feature enhancement module. The specific process is as follows: Step 3.1: Input feature f i Enter three different branches respectively and get three different outputs F1, F2, F3. The calculation process can be described as: in, They represent convolution operations with kernel sizes of 3×3, 1×1, 1×3, and 3×1, respectively. Represents a dilated convolution operation with a kernel size of 3×3 and a dilation rate of 5; Step 3.2: After the input features pass through three different branches, they are merged and concatenated with the residual connection after 1×1 convolution to obtain the final output f out , and its calculation process can be described as: Among them, F cat Indicates the concatenation operation of the feature map. Represents a standard convolution operation with a kernel size of 1×1.
8. The cross-modal person re-identification method based on coal mine scene according to claim 2 is characterized by: In step 4, the orthogonal projection loss function, the cross entropy loss function, the triple loss function, the center-guided pair mining loss function, and the KL divergence loss function are used to jointly constrain the model training process. Encoder and The features after the encoder are optimized by the KL divergence loss function. Encoder and The features after the encoder are optimized through the triplet loss function, and the features after the visible light-infrared shared path are jointly optimized through the center-guided mining loss function, the cross entropy loss function and the orthogonal projection loss function.
9. The cross-modal person re-identification method based on coal mine scene according to claim 8 is characterized in that: The calculation process of the orthogonal projection loss function is: First, define the calculation formula of the cosine similarity operator of two vectors, as shown in the following formula: Among them, ‖g‖2 represents the norm calculation process of the norm l2 operator; Secondly, the clustering calculation process of samples of the same category and the orthogonality calculation process of samples of different categories are defined as follows: Among them, <x i ,y i > is an input-output pair and calculates the cosine similarity operator of the two vectors, represents the output of the middle layer of the network, and B represents the size of the mini-batch; Finally, the total orthogonal projection loss function is expressed as: L ort =(1-d)+h· Among them, |·| is the absolute value operator.
10. The cross-modal person re-identification method based on coal mine scene according to claim 8, characterized in that: The cross entropy loss function is used for identity recognition in the pedestrian retrieval process. The cross-modal pedestrian re-identification task is regarded as an image classification task, and its calculation process can be expressed as: Among them, n represents the number of training samples, x i Represents the input of pedestrian image, y i represents its corresponding identity label, p(y i |x i ) represents the predicted probability of the input image and its corresponding identity label being identified after classification by the Softmax function; The triplet loss function is used to optimize the triplet relationship between the unified identity and different identities, ensuring that the feature distance of the same pedestrian image is smaller than the feature distance between different pedestrians. The cross-modal pedestrian re-identification task is regarded as an image retrieval task, and its calculation process can be expressed as: in, and Represent anchor samples, positive samples and negative samples respectively. and The IDs of people in the same group, and Then use the IDs of different pedestrians respectively, F(·) represents the feature extraction function, It represents the process of solving the Euclidean distance between positive and negative samples. The most difficult positive sample is the sample that belongs to the same category as the anchor sample and is the farthest away. The most difficult negative sample is the sample that belongs to a different category from the anchor sample and is closest to it. P represents the number of pedestrians in a single batch, K represents the number of pictures corresponding to each pedestrian, and α is the edge parameter. The center-guided pair mining loss function is used to constrain the generated embedding information, and the generated embedding is constrained by the following three attribute constraints, which increases the distance between the generated embedding and the original embedding, shortens the distance between the generated embedding of the visible light image and the original embedding of the infrared image, and shortens the distance between the generated embedding of the original image and the original embedding of the visible light image, and ensures that the intra-class distance is smaller than the inter-class distance, where α is a balance parameter, and its calculation process is expressed as: Similarly, for the control of the embedded class center generated by the infrared image, the calculation process can be expressed as: The final center-guided pair mining loss function can be expressed as: The calculation process of the KL divergence loss function is: Among them, p n,m For {f n ,f m } pairing probability, q n,m The constant ε is set to 1e for its true matching probability. -8 .
Citation Information
Patent Citations
Visible light infrared pedestrian re-identification method based on multi-modal relation aggregation
CN114511878A
Novel multi-modal fusion pedestrian re-identification algorithm
CN114694089A